Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
TL;DR
ML Systems Engineer (AI/RL Infrastructure): Building and maintaining distributed training infrastructure for SFT, pretraining, and RL workloads with an accent on GPU performance and scalability. Focus on implementing parallelism strategies, diagnosing distributed training failures, and optimizing GPU utilization for frontier models.
Location: Palo Alto, California, United States. Applicants must be authorized to work in the US.
Salary: $195,200 - $262,200 USD
Company
Nebius is building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment.
What you will do
- Build and maintain distributed training infrastructure for SFT, continued pretraining, and RL workloads.
- Integrate and extend frameworks such as Megatron-LM, DeepSpeed, PyTorch FSDP/DTensor, and Ray.
- Implement and debug complex parallelism strategies including tensor, pipeline, sequence, expert, and data parallelism.
- Develop reliable rollout, reward model serving, and experiment orchestration components for RL training.
- Profile and improve GPU utilization, communication efficiency, and training throughput.
- Collaborate with research scientists to transform algorithmic recipes into scalable, debuggable systems.
Requirements
- Strong Python and PyTorch engineering skills.
- Hands-on experience with distributed model training and large-scale ML systems or GPU cluster workloads.
- Deep understanding of transformer training bottlenecks, memory pressure, and communication overhead.
- Proven experience debugging production training jobs across multiple GPUs or nodes.
- Ability to quantitatively reason about throughput, utilization, memory, and cost.
- Must have valid US work authorization.
Nice to have
- Experience with Slurm, Kubernetes, or large-scale internal training platforms.
- Familiarity with RL infrastructure frameworks like verl, OpenRLHF, or custom PPO/GRPO systems.
- Knowledge of NCCL, CUDA, Triton, InfiniBand, RDMA, or H100/B200 clusters.
- Open-source contributions to distributed training or RL infrastructure projects.
Culture & Benefits
- 100% company-paid medical, dental, and vision coverage for employees and families.
- 401(k) plan with up to 4% company match and immediate vesting.
- Generous parental leave (20 weeks for primary, 12 weeks for secondary caregivers).
- Remote work reimbursement for mobile and internet.
- Company-paid short-term, long-term, and life insurance.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →