Назад
Company hidden
4 дня назад

Data Engineer (AI)

Формат работы
remote (только Argentina/Brazil/Colombia)
Тип работы
fulltime
Грейд
senior
Английский
c1
Страна
US/Argentina/Chile +4 еще
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Data Engineer (AI) (Python/SQL, Spark, Kafka, Airflow): Building production ingestion and transformation pipelines, warehouse and lakehouse systems, and retrieval infrastructure for AI and RAG applications with an accent on data quality, scalability, security, and cost control. Focus on designing reliable batch and streaming workflows, managing schema evolution and sensitive data, and operating vector retrieval systems across cloud environments.

Location: Fully remote across Argentina, Brazil, Colombia, Mexico, Chile, and Peru, aligned to the client's working day and United States time zones.

Company

hirify.global is a San Francisco-based software development company that builds and operates production AI systems and provides nearshore AI engineering teams.

What you will do

  • Build and operate batch and streaming ingestion and transformation pipelines using Spark, Kafka, dbt, and Airflow.
  • Design warehouses and lakehouses on Snowflake, BigQuery, Redshift, or Databricks, including partitioning, file layout, performance, and cost controls.
  • Develop chunking, embedding, indexing, and retrieval pipelines for RAG systems using pgvector, Pinecone, Qdrant, or Azure AI Search.
  • Establish data quality contracts with tests, lineage, expectations, alerting, and reliable handling of schema evolution, backfills, and late-arriving data.
  • Implement sensitive-data controls including PII classification, masking, access management, retention, deletion, and audit trails.
  • Deploy and operate containerized data systems on Azure or AWS with CI/CD, orchestration, observability, and explicit runtime budgets while working in client environments.

Requirements

  • 5+ years of experience building and operating production data pipelines with Python and SQL.
  • Deep expertise in data warehouse or lakehouse design, dimensional modeling, incremental processing, and cost-performance trade-offs.
  • Production-scale distributed processing with Spark, Kafka, Flink, or equivalent, plus orchestration with Airflow, Dagster, or Prefect.
  • Experience with transformation under version control, testing, lineage, cloud deployment, Docker, CI/CD, and infrastructure as code.
  • Active use of AI-assisted coding tools such as Claude Code, Cursor, or GitHub Copilot in production delivery.
  • Clear written and spoken English at C1 level or above, with the ability to explain technical trade-offs directly to clients; a bachelor's degree or equivalent professional experience.

Nice to have

  • Experience with vector and retrieval infrastructure, streaming or high-throughput workloads, and managed cloud data services.
  • Experience with Jupyter, Google Colab, or similar notebooks.
  • Delivery under SOC 2, HIPAA, or another compliance regime.
  • Open-source contributions, technical writing, or active participation in the data engineering community.

Culture & Benefits

  • AI-assisted coding tools are part of the standard engineering toolchain.
  • Automated codebase audits evaluate security, cost, and architecture findings.
  • Vendor-neutral work across OpenAI, Anthropic, and open-weight models, with access to hirify.global's Valkyrie production layer.
  • Paid time off, U.S. holidays, AI training, mentored career development, and profit sharing.
  • USD remuneration and opportunities for open-source work, community teaching, and philanthropy.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →