Назад
Company hidden
9 часов назад

ML Training Infrastructure Engineer (AI)

220 000 - 320 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
ML Training Infrastructure Engineer (AI): Building end-to-end infrastructure that turns a multi-cloud GPU fleet into a reliable training engine for embodied AI and multimodal robotics models with an accent on distributed training, high-throughput data pipelines, and production inference. Focus on optimizing GPU utilization, implementing reproducible scheduling and failure recovery, and compiling low-latency models for real-time robot control.

Location: Redwood City, California, United States; on-site

Salary: $220,000–$320,000 base salary per year, plus equity

Company

Builds general-purpose robots powered by an embodied AI foundation model and deployed across multiple commercial industries.

What you will do

  • Architect and own large-scale, multi-cloud GPU training infrastructure for massive multimodal models.
  • Implement distributed training techniques including sharding, activation checkpointing, mixed precision, and memory optimization with ZeRO and FSDP.
  • Build research codebases and Kubernetes/SLURM scheduling systems for fast iteration, automated retries, and failure recovery.
  • Design high-throughput pipelines for terabytes of multimodal robot data, including video, proprioception, and 3D signals.
  • Develop low-latency inference pipelines for real-time robot control using quantization, distillation, TensorRT, and Triton.
  • Profile GPU utilization, I/O bottlenecks, memory fragmentation, and inter-node communication across the compute fleet.

Requirements

  • 7+ years of engineering experience, including leadership of technical projects in HPC or ML infrastructure.
  • Deep experience with PyTorch and distributed training frameworks such as DeepSpeed and Accelerate.
  • Hands-on experience managing cloud GPU environments on GCP or AWS and using Kubernetes.
  • Understanding of distributed systems, race conditions, memory management, NCCL, and inter-node communication.
  • Ability to design, build, and operate infrastructure end to end.
  • Availability to work on-site in Redwood City, California, United States.

Nice to have

  • Experience with robotics data formats such as MCAP and Protobuf or multimodal models such as VLAs.
  • Experience with custom Triton kernels, compilers, or runtime optimization.
  • Experience as a founding or early-stage infrastructure hire.

Culture & Benefits

  • Work on robotics technology designed for real-world commercial applications.
  • Collaborate with experienced researchers and engineers from major technology companies.
  • Receive equity in addition to the base salary.
  • Work in an environment emphasizing technical rigor, mutual respect, problem-solving, and grit.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →