Назад
Company hidden
5 часов Π½Π°Π·Π°Π΄

Machine Learning Infrastructure Engineer

Π€ΠΎΡ€ΠΌΠ°Ρ‚ Ρ€Π°Π±ΠΎΡ‚Ρ‹
onsite
Π’ΠΈΠΏ Ρ€Π°Π±ΠΎΡ‚Ρ‹
fulltime
Π“Ρ€Π΅ΠΉΠ΄
senior
Английский
b2
Π‘Ρ‚Ρ€Π°Π½Π°
US
Вакансия ΠΈΠ· списка Hirify.GlobalВакансия ΠΈΠ· Hirify Global, списка ΠΌΠ΅ΠΆΠ΄ΡƒΠ½Π°Ρ€ΠΎΠ΄Π½Ρ‹Ρ… tech-ΠΊΠΎΠΌΠΏΠ°Π½ΠΈΠΉ
Для мэтча ΠΈ ΠΎΡ‚ΠΊΠ»ΠΈΠΊΠ° Π½ΡƒΠΆΠ΅Π½ Plus

ΠœΡΡ‚Ρ‡ & Π‘ΠΎΠΏΡ€ΠΎΠ²ΠΎΠ΄

Для мэтча с этой вакансиСй Π½ΡƒΠΆΠ΅Π½ Plus

ОписаниС вакансии

ВСкст:
/
TL;DR
Machine Learning Infrastructure Engineer (AI): Build and optimize large-scale training infrastructure and core model code with an accent on scalable JAX training pipelines and efficient GPU/TPU compute management. Focus on designing distributed training systems, performance optimization, and enabling rapid iteration for research experiments.

Location

Location: San Francisco, onsite

What you will do

  • Own and maintain large-scale model training infrastructure including scheduling, job management, checkpointing, and metrics/logging.
  • Scale distributed JAX-based training across TPU and GPU clusters.
  • Optimize performance by profiling memory usage, device utilization, throughput, and synchronization.
  • Build abstractions for launching, monitoring, debugging, and reproducing experiments.
  • Manage cloud-based GPU/TPU compute resources efficiently while controlling costs.
  • Collaborate with researchers to translate needs into infrastructure capabilities and best practices.
  • Contribute to core JAX training code supporting new architectures and evaluation metrics.

Requirements

  • Strong software engineering fundamentals and experience building ML training infrastructure or internal platforms.
  • Hands-on experience with large-scale training in JAX (preferred) or PyTorch.
  • Familiarity with distributed training, multi-host setups, data loaders, and evaluation pipelines.
  • Experience managing training workloads on cloud platforms such as SLURM, Kubernetes, GCP TPU/GKE, or AWS.
  • Ability to debug and optimize performance bottlenecks across the training stack.
  • Strong cross-functional communication and ownership mindset.

Nice to have

  • Deep ML systems background including training compilers, runtime optimization, and custom kernels.
  • Experience with GPU/TPU performance tuning and operating close to hardware.
  • Background in robotics, multimodal models, or large-scale foundation models.
  • Experience designing abstractions balancing researcher flexibility with system reliability.

Culture & Benefits

  • Consideration of qualified applicants with arrest and conviction records pursuant to San Francisco Fair Chance Ordinance.

Π‘ΡƒΠ΄ΡŒΡ‚Π΅ остороТны: Ссли Ρ€Π°Π±ΠΎΡ‚ΠΎΠ΄Π°Ρ‚Π΅Π»ΡŒ просит Π²ΠΎΠΉΡ‚ΠΈ Π² ΠΈΡ… систСму, ΠΈΡΠΏΠΎΠ»ΡŒΠ·ΡƒΡ iCloud/Google, ΠΏΡ€ΠΈΡΠ»Π°Ρ‚ΡŒ ΠΊΠΎΠ΄/ΠΏΠ°Ρ€ΠΎΠ»ΡŒ, Π·Π°ΠΏΡƒΡΡ‚ΠΈΡ‚ΡŒ ΠΊΠΎΠ΄/ПО, Π½Π΅ Π΄Π΅Π»Π°ΠΉΡ‚Π΅ этого - это мошСнники. ΠžΠ±ΡΠ·Π°Ρ‚Π΅Π»ΡŒΠ½ΠΎ ΠΆΠΌΠΈΡ‚Π΅ "ΠŸΠΎΠΆΠ°Π»ΠΎΠ²Π°Ρ‚ΡŒΡΡ" ΠΈΠ»ΠΈ ΠΏΠΈΡˆΠΈΡ‚Π΅ Π² ΠΏΠΎΠ΄Π΄Π΅Ρ€ΠΆΠΊΡƒ. ΠŸΠΎΠ΄Ρ€ΠΎΠ±Π½Π΅Π΅ Π² Π³Π°ΠΉΠ΄Π΅ β†’