9 часов назад
ML Training Infrastructure Engineer (AI)
220 000 - 320 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
ML Training Infrastructure Engineer (AI): Building end-to-end infrastructure that turns a multi-cloud GPU fleet into a reliable training engine for embodied AI and multimodal robotics models with an accent on distributed training, high-throughput data pipelines, and production inference. Focus on optimizing GPU utilization, implementing reproducible scheduling and failure recovery, and compiling low-latency models for real-time robot control.
Location: Redwood City, California, United States; on-site
Salary: $220,000–$320,000 base salary per year, plus equity
Company
Builds general-purpose robots powered by an embodied AI foundation model and deployed across multiple commercial industries.
What you will do
- Architect and own large-scale, multi-cloud GPU training infrastructure for massive multimodal models.
- Implement distributed training techniques including sharding, activation checkpointing, mixed precision, and memory optimization with ZeRO and FSDP.
- Build research codebases and Kubernetes/SLURM scheduling systems for fast iteration, automated retries, and failure recovery.
- Design high-throughput pipelines for terabytes of multimodal robot data, including video, proprioception, and 3D signals.
- Develop low-latency inference pipelines for real-time robot control using quantization, distillation, TensorRT, and Triton.
- Profile GPU utilization, I/O bottlenecks, memory fragmentation, and inter-node communication across the compute fleet.
Requirements
- 7+ years of engineering experience, including leadership of technical projects in HPC or ML infrastructure.
- Deep experience with PyTorch and distributed training frameworks such as DeepSpeed and Accelerate.
- Hands-on experience managing cloud GPU environments on GCP or AWS and using Kubernetes.
- Understanding of distributed systems, race conditions, memory management, NCCL, and inter-node communication.
- Ability to design, build, and operate infrastructure end to end.
- Availability to work on-site in Redwood City, California, United States.
Nice to have
- Experience with robotics data formats such as MCAP and Protobuf or multimodal models such as VLAs.
- Experience with custom Triton kernels, compilers, or runtime optimization.
- Experience as a founding or early-stage infrastructure hire.
Culture & Benefits
- Work on robotics technology designed for real-world commercial applications.
- Collaborate with experienced researchers and engineers from major technology companies.
- Receive equity in addition to the base salary.
- Work in an environment emphasizing technical rigor, mutual respect, problem-solving, and grit.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
5 часов назад
ML Engineer (AI Robotics)
200 000 - 350 000$
7 часов назад
Senior/Staff AI Algorithms Engineer (Robotics)
170 000 - 225 000$
1 день назад
Forward Deployed Engineer (Post-Sales) (AI)
230 000 - 300 000$
4 часа назад
AI Engineer (MedTech)
120 000 - 210 000$
5 часов назад
Technical Staff
200 000 - 350 000$
8 часов назад
AI Engineer
159 000 - 195 000$