5 часов назад
ML Infra Engineer (Data Systems) (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
ML Infra Engineer (Data Systems) (AI): Building and operating data infrastructure for large-scale robot learning with an accent on distributed systems, multimodal data pipelines, storage, and training-time performance. Focus on processing petabyte-scale datasets, optimizing dataloaders and data movement, and implementing observability and validation for reliable machine learning workflows.
Location: San Francisco, United States; on-site
Company
develops foundation models and learning algorithms for robots and physically actuated devices.
What you will do
- Design and build high-throughput pipelines for validating, transforming, and featurizing raw multimodal data.
- Operate large-scale batch and streaming workflows over massive datasets.
- Design object storage layouts, metadata systems, file formats, and efficient data access patterns.
- Build data lifecycle systems for backfills, dataset rebuilds, garbage collection, and large-scale transformations.
- Optimize dataloaders, sharding, prefetching, caching, and throughput to reduce the time from data arrival to model training.
- Implement metadata indexing, petabyte-scale data movement, observability, validation, and guardrails while collaborating with researchers, engineers, and roboticists.
Requirements
- Strong software engineering fundamentals.
- Experience building distributed systems or large-scale data pipelines.
- Ability to reason about performance, memory, I/O, and storage efficiency.
- Familiarity with batch and/or streaming processing systems.
- Experience with object storage systems and data format tradeoffs.
- Ownership of designing, building, operating, and iterating on systems end to end.
Nice to have
- Experience with large machine learning training pipelines or dataloading systems.
- Knowledge of columnar or custom data formats.
- Experience with ClickHouse, Ray, Flink, Spark, or similar systems.
- Hands-on experience operating petabyte-scale datasets.
- Experience debugging and fixing performance bottlenecks in data-heavy systems.
Culture & Benefits
- Work on-site within an infrastructure organization supporting large-scale learning.
- Collaborate closely with researchers, engineers, and roboticists on fast-moving projects.
- Build and operate systems that prioritize performance, correctness, reliability, and operational ownership.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
6 часов назад
Distributed Training Infrastructure Engineer (AI)
6 часов назад
Customer Engineer (ML/AI)
170 000 - 199 000$
6 часов назад
Applied AI Engineer (Robotics)
250 000 - 300 000$
6 часов назад
Platform Engineer (AI)
Cognition
6 дней назад
Software Engineer (AI)
260 000 - 300 000$
5 часов назад