5 ΡΠ°ΡΠΎΠ² Π½Π°Π·Π°Π΄
Machine Learning Infrastructure Engineer
ΠΡΡΡ & Π‘ΠΎΠΏΡΠΎΠ²ΠΎΠ΄
ΠΠ»Ρ ΠΌΡΡΡΠ° Ρ ΡΡΠΎΠΉ Π²Π°ΠΊΠ°Π½ΡΠΈΠ΅ΠΉ Π½ΡΠΆΠ΅Π½ Plus
ΠΠΏΠΈΡΠ°Π½ΠΈΠ΅ Π²Π°ΠΊΠ°Π½ΡΠΈΠΈ
Π’Π΅ΠΊΡΡ:
TL;DR
Machine Learning Infrastructure Engineer (AI): Build and optimize large-scale training infrastructure and core model code with an accent on scalable JAX training pipelines and efficient GPU/TPU compute management. Focus on designing distributed training systems, performance optimization, and enabling rapid iteration for research experiments.
Location
Location: San Francisco, onsite
What you will do
- Own and maintain large-scale model training infrastructure including scheduling, job management, checkpointing, and metrics/logging.
- Scale distributed JAX-based training across TPU and GPU clusters.
- Optimize performance by profiling memory usage, device utilization, throughput, and synchronization.
- Build abstractions for launching, monitoring, debugging, and reproducing experiments.
- Manage cloud-based GPU/TPU compute resources efficiently while controlling costs.
- Collaborate with researchers to translate needs into infrastructure capabilities and best practices.
- Contribute to core JAX training code supporting new architectures and evaluation metrics.
Requirements
- Strong software engineering fundamentals and experience building ML training infrastructure or internal platforms.
- Hands-on experience with large-scale training in JAX (preferred) or PyTorch.
- Familiarity with distributed training, multi-host setups, data loaders, and evaluation pipelines.
- Experience managing training workloads on cloud platforms such as SLURM, Kubernetes, GCP TPU/GKE, or AWS.
- Ability to debug and optimize performance bottlenecks across the training stack.
- Strong cross-functional communication and ownership mindset.
Nice to have
- Deep ML systems background including training compilers, runtime optimization, and custom kernels.
- Experience with GPU/TPU performance tuning and operating close to hardware.
- Background in robotics, multimodal models, or large-scale foundation models.
- Experience designing abstractions balancing researcher flexibility with system reliability.
Culture & Benefits
- Consideration of qualified applicants with arrest and conviction records pursuant to San Francisco Fair Chance Ordinance.
ΠΡΠ΄ΡΡΠ΅ ΠΎΡΡΠΎΡΠΎΠΆΠ½Ρ: Π΅ΡΠ»ΠΈ ΡΠ°Π±ΠΎΡΠΎΠ΄Π°ΡΠ΅Π»Ρ ΠΏΡΠΎΡΠΈΡ Π²ΠΎΠΉΡΠΈ Π² ΠΈΡ ΡΠΈΡΡΠ΅ΠΌΡ, ΠΈΡΠΏΠΎΠ»ΡΠ·ΡΡ iCloud/Google, ΠΏΡΠΈΡΠ»Π°ΡΡ ΠΊΠΎΠ΄/ΠΏΠ°ΡΠΎΠ»Ρ, Π·Π°ΠΏΡΡΡΠΈΡΡ ΠΊΠΎΠ΄/ΠΠ, Π½Π΅ Π΄Π΅Π»Π°ΠΉΡΠ΅ ΡΡΠΎΠ³ΠΎ - ΡΡΠΎ ΠΌΠΎΡΠ΅Π½Π½ΠΈΠΊΠΈ. ΠΠ±ΡΠ·Π°ΡΠ΅Π»ΡΠ½ΠΎ ΠΆΠΌΠΈΡΠ΅ "ΠΠΎΠΆΠ°Π»ΠΎΠ²Π°ΡΡΡΡ" ΠΈΠ»ΠΈ ΠΏΠΈΡΠΈΡΠ΅ Π² ΠΏΠΎΠ΄Π΄Π΅ΡΠΆΠΊΡ. ΠΠΎΠ΄ΡΠΎΠ±Π½Π΅Π΅ Π² Π³Π°ΠΉΠ΄Π΅ β
ΠΠΎΡ ΠΎΠΆΠΈΠ΅ Π²Π°ΠΊΠ°Π½ΡΠΈΠΈ
Windsurf
6 Π΄Π½Π΅ΠΉ Π½Π°Π·Π°Π΄
Research Engineer (AI)
6 ΡΠ°ΡΠΎΠ² Π½Π°Π·Π°Π΄
Software Engineer (AI Inference)
175Β 000 - 275Β 000$
6 ΡΠ°ΡΠΎΠ² Π½Π°Π·Π°Π΄
Distributed Training Infrastructure Engineer (AI)
5 Π΄Π½Π΅ΠΉ Π½Π°Π·Π°Π΄
ML Infrastructure Engineer (AI)
14Β 167 - 25Β 000$
6 ΡΠ°ΡΠΎΠ² Π½Π°Π·Π°Π΄
Research Scientist (Embodied AI)
250Β 000 - 350Β 000$
5 ΡΠ°ΡΠΎΠ² Π½Π°Π·Π°Π΄