Назад
7 дней назад

ML Systems Engineer (AI/RL Infrastructure)

195 200 - 262 200$
Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify RU Global, списка компаний с восточно-европейскими корнями
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/

TL;DR

ML Systems Engineer (AI/RL Infrastructure): Building and maintaining distributed training infrastructure for SFT, pretraining, and RL workloads with an accent on GPU performance and scalability. Focus on implementing parallelism strategies, diagnosing distributed training failures, and optimizing GPU utilization for frontier models.

Location: Palo Alto, California, United States. Applicants must be authorized to work in the US.

Salary: $195,200 - $262,200 USD

Company

Nebius is building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment.

What you will do

  • Build and maintain distributed training infrastructure for SFT, continued pretraining, and RL workloads.
  • Integrate and extend frameworks such as Megatron-LM, DeepSpeed, PyTorch FSDP/DTensor, and Ray.
  • Implement and debug complex parallelism strategies including tensor, pipeline, sequence, expert, and data parallelism.
  • Develop reliable rollout, reward model serving, and experiment orchestration components for RL training.
  • Profile and improve GPU utilization, communication efficiency, and training throughput.
  • Collaborate with research scientists to transform algorithmic recipes into scalable, debuggable systems.

Requirements

  • Strong Python and PyTorch engineering skills.
  • Hands-on experience with distributed model training and large-scale ML systems or GPU cluster workloads.
  • Deep understanding of transformer training bottlenecks, memory pressure, and communication overhead.
  • Proven experience debugging production training jobs across multiple GPUs or nodes.
  • Ability to quantitatively reason about throughput, utilization, memory, and cost.
  • Must have valid US work authorization.

Nice to have

  • Experience with Slurm, Kubernetes, or large-scale internal training platforms.
  • Familiarity with RL infrastructure frameworks like verl, OpenRLHF, or custom PPO/GRPO systems.
  • Knowledge of NCCL, CUDA, Triton, InfiniBand, RDMA, or H100/B200 clusters.
  • Open-source contributions to distributed training or RL infrastructure projects.

Culture & Benefits

  • 100% company-paid medical, dental, and vision coverage for employees and families.
  • 401(k) plan with up to 4% company match and immediate vesting.
  • Generous parental leave (20 weeks for primary, 12 weeks for secondary caregivers).
  • Remote work reimbursement for mobile and internet.
  • Company-paid short-term, long-term, and life insurance.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →