Назад
Company hidden
5 часов назад

Distributed Training Infrastructure Engineer (AI)

Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Distributed Training Infrastructure Engineer (AI): Design and build scalable distributed training infrastructure for GPU clusters powering next-generation AI foundation models with an accent on runtime performance, reliability, and optimization of parallelism and data pipelines. Focus on building core systems from scratch, solving distributed systems complexity, and improving training stability and efficiency.

Location

Location: San Francisco, Boston, or Remote (Hybrid) with preference for San Francisco and Boston locations in the United States.

Company

hirify.global is a startup spun out of MIT CSAIL building general-purpose AI systems optimized for diverse deployment targets including data centers and on-device hardware.

What you will do

  • Design and build core distributed training systems for large-scale GPU clusters.
  • Implement and optimize parallelism and sharding strategies for evolving AI model architectures.
  • Improve distributed training efficiency through topology-aware collectives and straggler mitigation.
  • Develop data loading systems to eliminate I/O bottlenecks for multimodal datasets.
  • Create checkpointing mechanisms balancing memory and recovery needs.
  • Build monitoring, profiling, and debugging tools to ensure training stability and performance.

Requirements

  • Hands-on experience with distributed training infrastructure (PyTorch Distributed DDP/FSDP, DeepSpeed ZeRO, Megatron-LM TP/PP).
  • Experience diagnosing performance bottlenecks and failure modes including profiling and NCCL/collectives issues.
  • Understanding of hardware accelerators and networking topologies.
  • Experience optimizing data pipelines for machine learning workloads.

Nice to have

  • Experience with Mixture of Experts (MoE) training.
  • Experience with large-scale distributed training (100+ GPUs).
  • Open-source contributions to training infrastructure projects.

Culture & Benefits

  • Competitive base salary with equity in a unicorn-stage startup.
  • 100% paid medical, dental, and vision insurance for employees and dependents.
  • 401(k) matching up to 4% of base pay.
  • Unlimited PTO plus company-wide Refill Days.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →