обновлено 59 минут назад
Infrastructure Reliability Engineer (AI Data Infrastructure)
125 000 - 170 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Infrastructure Reliability Engineer (AI Data Infrastructure): Build and operate large-scale data systems for AI training, evaluation, and continual improvement with an accent on multimodal ingestion, data quality, lineage, and high-throughput delivery. Focus on designing petabyte-scale pipelines, maximizing GPU utilization, enforcing privacy and provenance controls, and optimizing storage cost and performance.
Location: 100% remote within the United States
Salary: $125,000–$170,000 annually
Company
is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.
What you will do
- Design and operate large-scale data pipelines supporting AI training, evaluation, and continual improvement.
- Build ingestion, cleaning, deduplication, filtering, and quality-assurance systems for text, image, audio, video, and structured data.
- Develop dataset versioning, lineage, provenance, evaluation, and reproducibility systems.
- Build high-throughput data loading systems and storage architectures that balance GPU utilization, cost, throughput, and latency.
- Implement labeling, active learning, human-in-the-loop improvement, privacy, redaction, and consent workflows.
- Drive observability, documentation, cross-functional alignment, and cost and performance optimization across AI data infrastructure.
Requirements
- Bachelor’s or Master’s degree in Computer Science or a related field.
- 6+ years of data engineering experience, including significant work with ML or AI workloads.
- Strong proficiency in Python and at least one JVM or systems language.
- Deep experience with Spark, Ray, or Beam, plus hands-on operation of petabyte-scale storage and pipeline systems.
- Strong understanding of distributed systems, data modeling, storage formats, dataset versioning, lineage, and ML reproducibility.
- Experience with testing, CI/CD, code review, communication, and cross-functional collaboration.
Nice to have
- Experience with large-scale multimodal datasets and frontier model training pipelines.
- Familiarity with data quality tooling and dataset evaluation methodology.
- Exposure to privacy-preserving systems and regulated data handling.
- Open-source contributions to data infrastructure projects.
Culture & Benefits
- Full-time direct W2 employment.
- Fully remote work within the United States.
- Career growth opportunity within an established organization.
- Equal employment opportunity and a workplace free from discrimination and harassment.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →