Назад
обновлено 8 дней назад

Post-Training Researcher (AI)

Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Post-Training Researcher (AI): Developing post-training recipes, evaluations, and alignment methods for AI agents such as Devin with an accent on reinforcement learning, preference modeling, and large-scale distributed training. Focus on designing trustworthy evaluations, debugging unexpected training behavior, measuring scaling effects across data and compute, and improving long-horizon agent reasoning and collaboration.

Location: San Francisco, United States; on-site

Company

Building Devin, an AI software engineer, and advancing AI systems that reason about real-world tasks.

What you will do

  • Develop and iterate on post-training recipes covering datasets, training stages, and hyperparameters.
  • Design evaluations that measure meaningful model and production performance while maintaining evaluation integrity.
  • Investigate unexpected training results and identify the underlying causes of model behavior.
  • Apply and advance RLHF, RLAIF, preference modeling, reward learning, and constitutional approaches for agent alignment.
  • Measure scaling behavior across data and compute, and develop new methodologies when existing approaches reach their limits.
  • Work across research and engineering to move prototypes into real deployment.

Requirements

  • Demonstrated work advancing ML systems through post-training, alignment, or related methods.
  • Strong foundations in probability, statistics, and machine learning theory.
  • Original contributions through publications, open-source work, or equivalent industry results.
  • Experience with large-scale distributed training and debugging.
  • Systems-level understanding of the interaction between training pipelines, data, and evaluation.
  • Ability to work effectively in ambiguous, fast-moving research environments.

Nice to have

  • PhD or another credential demonstrating research capability; credentials are considered alongside demonstrated ability.

Culture & Benefits

  • Small, highly selective team where research and product development move together.
  • Rapid path from prototypes to real deployment.
  • Large compute allocations with training jobs running across thousands of GPUs.
  • Environment emphasizing speed, autonomy, technical depth, and minimal process overhead.
  • Resources for operating at frontier scale from the start.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →