Назад
Company hidden
9 часов назад

ML Infrastructure Engineer (AI)

Формат работы
onsite
Тип работы
fulltime
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
ML Infrastructure Engineer (AI): Building and optimizing large-scale training systems for foundation models with an accent on JAX pipelines, distributed training, and GPU/TPU compute management. Focus on scaling model training from prototype to production and improving performance across the training stack.

Location: San Francisco (On-site)

Company

hirify.global is developing foundation models and learning algorithms to power general-purpose robots and physically-actuated devices.

What you will do

  • Design and maintain systems for large-scale model training, including scheduling, job management, and logging.
  • Scale JAX-based training across TPU and GPU clusters.
  • Profile and improve memory usage, device utilization, and throughput.
  • Build abstractions for launching, monitoring, and debugging experiments.
  • Partner with researchers to translate needs into infrastructure capabilities.
  • Evolve core training code to support new architectures and modalities.

Requirements

  • Strong software engineering fundamentals and experience building ML training infrastructure.
  • Hands-on experience with large-scale training in JAX or PyTorch.
  • Familiarity with distributed training, multi-host setups, and data pipelines.
  • Experience managing workloads on cloud platforms like Kubernetes, GCP, or AWS.
  • Ability to debug and optimize performance bottlenecks across the training stack.
  • Must be able to work on-site in San Francisco.

Nice to have

  • Deep ML systems background (compilers, runtime optimization, custom kernels).
  • Experience with GPU/TPU performance tuning.
  • Background in robotics or multimodal foundation models.

Culture & Benefits

  • Work at the intersection of ML, software engineering, and robotics.
  • High-leverage role impacting core modeling efforts.
  • Collaborative environment working closely with researchers and platform engineers.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →