4 часа назад
ML Infrastructure Engineer
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
ML Infrastructure Engineer (JAX/Distributed Training): Building and optimizing infrastructure for large-scale model training across GPU and TPU clusters with an accent on job orchestration, reusable JAX pipelines, and reliable experiment execution. Focus on profiling memory and device utilization, improving distributed synchronization, and translating research requirements into production-grade training systems.
Location: San Francisco, United States; on-site
Company
develops machine learning and robotics systems supported by large-scale model training.
What you will do
- Design, implement, and maintain infrastructure for large-scale model training and inference, including scheduling, job management, checkpointing, and metrics.
- Scale JAX-based distributed training across multi-host TPU and GPU clusters.
- Profile and optimize memory usage, device utilization, throughput, and distributed synchronization.
- Build abstractions for launching, monitoring, debugging, and reproducing experiments.
- Manage cloud GPU and TPU resources to improve utilization and control costs.
- Partner with researchers and evolve core JAX training code for new architectures, modalities, and evaluation metrics.
Requirements
- Strong software engineering fundamentals and experience building ML training infrastructure or internal platforms.
- Hands-on experience with large-scale training in JAX or PyTorch.
- Familiarity with distributed training, multi-host systems, data loaders, and evaluation pipelines.
- Experience managing workloads on cloud platforms and technologies such as SLURM, Kubernetes, GCP TPU/GKE, or AWS.
- Ability to debug and optimize performance bottlenecks across the training stack, including GPU or TPU performance.
- Strong cross-functional communication and ownership skills.
Nice to have
- Experience with training compilers, runtime optimization, or custom kernels.
- Background in robotics, multimodal models, or large-scale foundation models.
- Experience designing abstractions that balance researcher flexibility with system reliability.
Culture & Benefits
- Hands-on collaboration with research, data, platform, and model engineering teams.
- Work focused on making large-scale training reliable, reproducible, and fast.
- Close partnership with researchers to turn ideas into experiments and production training runs.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →