Назад
Company hidden
5 часов назад

ML Infra Engineer (Data Systems) (AI)

Формат работы
onsite
Тип работы
fulltime
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
ML Infra Engineer (Data Systems) (AI): Building and operating data infrastructure for large-scale robot learning with an accent on distributed systems, multimodal data pipelines, storage, and training-time performance. Focus on processing petabyte-scale datasets, optimizing dataloaders and data movement, and implementing observability and validation for reliable machine learning workflows.

Location: San Francisco, United States; on-site

Company

hirify.global develops foundation models and learning algorithms for robots and physically actuated devices.

What you will do

  • Design and build high-throughput pipelines for validating, transforming, and featurizing raw multimodal data.
  • Operate large-scale batch and streaming workflows over massive datasets.
  • Design object storage layouts, metadata systems, file formats, and efficient data access patterns.
  • Build data lifecycle systems for backfills, dataset rebuilds, garbage collection, and large-scale transformations.
  • Optimize dataloaders, sharding, prefetching, caching, and throughput to reduce the time from data arrival to model training.
  • Implement metadata indexing, petabyte-scale data movement, observability, validation, and guardrails while collaborating with researchers, engineers, and roboticists.

Requirements

  • Strong software engineering fundamentals.
  • Experience building distributed systems or large-scale data pipelines.
  • Ability to reason about performance, memory, I/O, and storage efficiency.
  • Familiarity with batch and/or streaming processing systems.
  • Experience with object storage systems and data format tradeoffs.
  • Ownership of designing, building, operating, and iterating on systems end to end.

Nice to have

  • Experience with large machine learning training pipelines or dataloading systems.
  • Knowledge of columnar or custom data formats.
  • Experience with ClickHouse, Ray, Flink, Spark, or similar systems.
  • Hands-on experience operating petabyte-scale datasets.
  • Experience debugging and fixing performance bottlenecks in data-heavy systems.

Culture & Benefits

  • Work on-site within an infrastructure organization supporting large-scale learning.
  • Collaborate closely with researchers, engineers, and roboticists on fast-moving projects.
  • Build and operate systems that prioritize performance, correctness, reliability, and operational ownership.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →