Назад
Company hidden
8 часов назад

AI Pipeline Engineer

100 000 - 150 000$
Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
AI Pipeline Engineer (Python/Spark/Ray): Building and operating petabyte-scale data pipelines for AI training, evaluation, and continual improvement with an accent on multimodal ingestion, data quality, lineage, and high-throughput delivery. Focus on dataset reproducibility, GPU-utilization optimization, privacy enforcement, and observability across distributed data systems.

Location: 100% remote within the United States

Salary: $100,000–$150,000 annually

Company

hirify.global is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.

What you will do

  • Design and operate large-scale data pipelines supporting AI training, evaluation, and continual improvement.
  • Build ingestion systems for text, image, audio, video, and structured data.
  • Implement data cleaning, deduplication, filtering, quality assurance, versioning, lineage, and provenance tracking at petabyte scale.
  • Develop high-throughput data loading, storage, caching, compression, and format strategies to improve training performance and GPU utilization.
  • Build labeling, active learning, human-in-the-loop, and evaluation dataset pipelines with integrity and contamination controls.
  • Implement privacy, redaction, consent, observability, and operational documentation across AI data systems while collaborating with ML researchers and engineers.

Requirements

  • Must be based in the United States and eligible to work there.
  • Bachelor’s or Master’s degree in Computer Science or a related field.
  • 6+ years of data engineering experience, including significant work supporting ML or AI workloads.
  • Strong Python proficiency and experience with at least one JVM or systems language.
  • Deep experience with Spark, Ray, or Beam and hands-on operation of petabyte-scale storage and pipeline systems.
  • Strong understanding of distributed systems, data modeling, storage formats, ML dataset reproducibility, testing, CI/CD, code review, and cross-functional collaboration.

Nice to have

  • Experience with large-scale multimodal datasets and frontier model training pipelines.
  • Familiarity with data quality tooling and dataset evaluation methodology.
  • Exposure to privacy-preserving data systems and regulated data handling.
  • Open-source contributions to data infrastructure projects.

Culture & Benefits

  • Full-time direct W-2 employment.
  • Career growth opportunities within an established technology organization.
  • Collaboration with ML researchers and engineers on modern AI infrastructure.
  • New H-1B visa petitions are not sponsored; U.S. citizens, Green Card holders, EAD holders, and H-1B transfer candidates may apply.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →