8 часов назад
AI Pipeline Engineer
100 000 - 150 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
AI Pipeline Engineer (Python/Spark/Ray): Building and operating petabyte-scale data pipelines for AI training, evaluation, and continual improvement with an accent on multimodal ingestion, data quality, lineage, and high-throughput delivery. Focus on dataset reproducibility, GPU-utilization optimization, privacy enforcement, and observability across distributed data systems.
Location: 100% remote within the United States
Salary: $100,000–$150,000 annually
Company
is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.
What you will do
- Design and operate large-scale data pipelines supporting AI training, evaluation, and continual improvement.
- Build ingestion systems for text, image, audio, video, and structured data.
- Implement data cleaning, deduplication, filtering, quality assurance, versioning, lineage, and provenance tracking at petabyte scale.
- Develop high-throughput data loading, storage, caching, compression, and format strategies to improve training performance and GPU utilization.
- Build labeling, active learning, human-in-the-loop, and evaluation dataset pipelines with integrity and contamination controls.
- Implement privacy, redaction, consent, observability, and operational documentation across AI data systems while collaborating with ML researchers and engineers.
Requirements
- Must be based in the United States and eligible to work there.
- Bachelor’s or Master’s degree in Computer Science or a related field.
- 6+ years of data engineering experience, including significant work supporting ML or AI workloads.
- Strong Python proficiency and experience with at least one JVM or systems language.
- Deep experience with Spark, Ray, or Beam and hands-on operation of petabyte-scale storage and pipeline systems.
- Strong understanding of distributed systems, data modeling, storage formats, ML dataset reproducibility, testing, CI/CD, code review, and cross-functional collaboration.
Nice to have
- Experience with large-scale multimodal datasets and frontier model training pipelines.
- Familiarity with data quality tooling and dataset evaluation methodology.
- Exposure to privacy-preserving data systems and regulated data handling.
- Open-source contributions to data infrastructure projects.
Culture & Benefits
- Full-time direct W-2 employment.
- Career growth opportunities within an established technology organization.
- Collaboration with ML researchers and engineers on modern AI infrastructure.
- New H-1B visa petitions are not sponsored; U.S. citizens, Green Card holders, EAD holders, and H-1B transfer candidates may apply.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
6 дней назад
Senior Data Engineer (Healthcare AI)
165 000 - 220 000$
2 дня назад
Senior People Data and AI Engineer
142 600 - 257 600$
Databricks
1 день назад
Staff Data Scientist (AI)
192 000 - 260 000$
3 дня назад
Senior Data Infrastructure Engineer (AI)
6 дней назад
Staff Machine Learning Engineer (AI)
118 400 - 171 000$
1 день назад
Senior Data Scientist (AI/ML)
124 000 - 329 200$