2 часа назад
AI Data Engineer
100 000 - 150 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
AI Data Engineer (AI/Data Infrastructure): Building and operating large-scale ingestion, transformation, quality, versioning, and delivery systems for multimodal AI training and evaluation pipelines with an accent on petabyte-scale data processing, lineage, privacy, and high-throughput loading. Focus on designing reproducible datasets, maximizing accelerator utilization, monitoring data quality and drift, and optimizing storage cost and performance.
Location: 100% remote within the United States
Salary: $100,000–$150,000 annually
Company
is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.
What you will do
- Design and operate large-scale data pipelines for AI training, evaluation, and continual improvement.
- Build ingestion, cleaning, deduplication, filtering, and quality assurance systems for text, image, audio, video, and structured data.
- Develop dataset versioning, lineage, provenance, evaluation, and contamination-control systems for reproducible ML workflows.
- Build high-throughput data loading and storage architectures that balance GPU utilization, cost, throughput, and latency.
- Implement labeling, active learning, human-in-the-loop, privacy, redaction, and consent-enforcement workflows.
- Collaborate with ML researchers and engineers while driving observability, documentation, and operational improvements across data systems.
Requirements
- Bachelor’s or Master’s degree in Computer Science or a related field.
- 10+ years of data engineering experience, including significant work supporting ML or AI workloads.
- Strong Python proficiency and proficiency in at least one JVM or systems language.
- Deep experience with Spark, Ray, or Beam, plus distributed systems, data modeling, and storage formats.
- Hands-on experience operating petabyte-scale storage and pipeline systems, including ML dataset versioning, lineage, and reproducibility.
- Strong software engineering practices, including testing, CI/CD, code review, communication, and cross-functional collaboration.
Nice to have
- Experience with large-scale multimodal datasets and frontier model training pipelines.
- Familiarity with data quality tooling and dataset evaluation methodology.
- Exposure to privacy-preserving systems and regulated data handling.
- Open-source contributions to data infrastructure projects.
Culture & Benefits
- Full-time direct W2 employment.
- Fully remote work within the United States.
- Career growth opportunities within an established technology consulting and software development organization.
Hiring process
- Submit a resume for consideration.
- U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply.
- New H-1B visa petition sponsorship is not available.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
6 дней назад
Senior Data Engineer (Healthcare AI)
165 000 - 220 000$
2 дня назад
Senior People Data and AI Engineer
142 600 - 257 600$
Databricks
1 день назад
Staff Data Scientist (AI)
192 000 - 260 000$
1 день назад
Senior Data Scientist (AI/ML)
124 000 - 329 200$
6 дней назад
Associate Data Scientist (AI/ML)
100 000 - 115 000$
6 дней назад
Staff Machine Learning Engineer (AI)
118 400 - 171 000$