22 часа назад
Senior Research Engineer, LLM Training & Post-Training (PyTorch)
165 000 - 310 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Research Engineer, LLM Training & Post-Training (PyTorch): Building and optimizing training, post-training, evaluation, and deployment systems for transformer-based language models with an accent on PyTorch, distributed multi-GPU training, and model quality. Focus on solving convergence and performance bottlenecks, improving memory efficiency and scalability, and translating research into production-ready AI systems.
Location: Hybrid with a minimum of 2 in-office days per week in San Francisco, Seattle, New York City, or London; fully remote work is considered for candidates outside the office hub locations.
Annual base salary: $165,000–$310,000 USD, plus discretionary bonus, equity, and benefits.
Company
builds an end-to-end platform for developing, training, and deploying AI systems, combining developer-first software with large-scale AI compute.
What you will do
- Design, build, and optimize training and post-training pipelines for large language models.
- Improve model quality through continued pretraining, supervised fine-tuning, preference optimization, reinforcement learning, reward modeling, evaluation, and experimentation.
- Build PyTorch-based training infrastructure, research tooling, and developer workflows.
- Optimize distributed multi-GPU training for throughput, memory efficiency, scalability, and GPU utilization.
- Investigate convergence, instability, communication overhead, and performance bottlenecks; design benchmarks and analyze model failure modes.
- Collaborate with research, infrastructure, platform engineering, and customers to develop production-ready AI systems and open-source improvements.
Requirements
- Significant experience training, fine-tuning, evaluating, and optimizing transformer-based language models with PyTorch.
- Experience with modern LLM training and post-training methods, including SFT, RLHF, DPO, PPO, GRPO, reward modeling, or similar approaches.
- Strong understanding of distributed training and multi-GPU systems, including performance, scalability, or efficiency optimization.
- Strong software engineering fundamentals and experience building production-quality Python software and research tooling.
- Experience designing experiments, evaluating model performance, and debugging complex training or optimization issues.
- Master’s degree, PhD, or equivalent industry experience in Machine Learning, AI, Computer Science, or a related field.
Nice to have
- Experience with DeepSpeed, FSDP, Megatron-LM, Hugging Face Transformers, TRL, PEFT, Lightning Fabric, or similar frameworks.
- Experience with CUDA, Triton, vLLM, SGLang, TensorRT, GPU optimization, mixed precision, or memory optimization.
- Open-source contributions, research publications, production AI platform experience, or startup experience.
Culture & Benefits
- Hybrid work model for office-based teams, flexible schedules, and occasional company and team offsites.
- Medical, dental, and vision coverage, meaningful equity, and retirement contributions, including 401(k) matching in the U.S. and pension contributions in the U.K.
- Unlimited PTO, company holidays, floating holidays, and a two-week company-wide winter break.
- Paid parental and family leave, wellness and work-from-home stipends, and an annual learning and development allowance.
- Four weeks of paid sabbatical leave after four years and complimentary meals at office hubs.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →