Назад
5 дней назад

Software Engineer (Reinforcement Learning)

142 800 - 274 800$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Software Engineer (Reinforcement Learning) (AI/LLM post-training): Building and scaling reinforcement learning environments, rollout systems, training pipelines, and data-generation workflows for frontier models and agents with an accent on reward modeling, evaluation, and distributed experimentation. Focus on improving reasoning, instruction following, tool use, coding, and agentic capabilities while designing reliable, observable, and reproducible systems for models used by millions of people.

Location: Mountain View, New York, or Redmond, United States

Salary: USD $142,800–$274,800 base pay per year across the U.S.; USD $188,000–$304,200 in the San Francisco Bay Area and New York City metropolitan area.

Company

Microsoft AI builds frontier models and systems that support Copilot and the broader Microsoft ecosystem.

What you will do

  • Set technical direction and lead multi-quarter initiatives for reinforcement learning and post-training of large language models and agents.
  • Build and scale RL environments, agent rollouts, training pipelines, and data-generation workflows.
  • Develop tasks, feedback signals, reward models, and datasets using human, model-generated, and real-world interaction data.
  • Design experiments to improve reasoning, instruction following, tool use, coding, and other agentic capabilities.
  • Define benchmarks, diagnose model failures, and evaluate core capabilities and real-world performance with evaluation teams.
  • Improve the reliability, observability, throughput, and reproducibility of large-scale RL systems while mentoring engineers and collaborating with research, infrastructure, safety, and product teams.

Requirements

  • Bachelor’s degree in Computer Science or a related technical field plus 6+ years of technical engineering experience, or equivalent experience.
  • Professional coding experience in C, C++, C#, Java, JavaScript, Python, or comparable languages.
  • Experience with reliable machine learning systems and reinforcement learning, reward modeling, preference optimization, model post-training, or related techniques for large-scale generative models.
  • Experience leading technically ambitious, multi-team machine learning initiatives from hypothesis and experiment design through measurable model or product impact.
  • Experience with RL environments, agent harnesses, simulators, distributed training or inference systems, and GPU-accelerated infrastructure.
  • Experience generating, curating, filtering, versioning, and evaluating large datasets and communicating complex experimental results to cross-functional partners.

Nice to have

  • Master’s degree in Computer Science or a related technical field with 8+ years of experience, or a bachelor’s degree with 12+ years of experience.
  • Familiarity with modern post-training methods and open-source machine learning frameworks.
  • Experience shaping RL, post-training, or evaluation roadmaps across multiple teams or model releases.

Culture & Benefits

  • Collaborative, fast-paced environment focused on rigorous experimentation and technical excellence.
  • Opportunity to ship model capabilities to millions of users through the Microsoft ecosystem.
  • Inclusive engineering culture grounded in knowledge sharing, mentorship, and end-to-end ownership.
  • Benefits and additional compensation may be available depending on role eligibility.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →