Назад
Company hidden
13 дней назад

Member of Technical Staff — RL Research (Experienced)

300 000 - 500 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Member of Technical Staff — RL Research (Experienced) (RL and post-training): Building Nuance’s large-scale RL and post-training stack for real-time audiovisual foundation models with an accent on rollout generation, reward modeling, policy optimization, and multimodal evaluation. Focus on designing distributed training infrastructure, improving interactive behavior and timing, and optimizing throughput, serving latency, GPU utilization, and research iteration speed.

Location: In-person in Seattle, five days a week

Salary: $300,000–$500,000 annual base salary, plus equity

Company

hirify.global is building photorealistic, real-time AI avatars with emotional intelligence through full-duplex audiovisual foundation models.

What you will do

  • Build Nuance’s RL and post-training stack from 0→1 and scale it from 1→10.
  • Develop rollout generation, policy optimization, reward and reference model serving, evaluation, checkpointing, observability, and debugging systems.
  • Design abstractions connecting research ideas with production-scale trainers, rollout workers, evaluators, data queues, experience buffers, and checkpoint promotion.
  • Develop feedback loops for turn-taking, interruption, timing, emotional response, audiovisual coherence, instruction following, and real-time interaction quality.
  • Optimize rollout throughput, serving latency, GPU utilization, policy updates, queueing, checkpoint overhead, and research iteration speed.

Requirements

  • Significant hands-on experience with RL, RLHF, RLAIF, post-training, alignment, or large-scale fine-tuning for foundation models.
  • Deep understanding of policy optimization, reward modeling, preference optimization, rejection sampling, KL control, evaluation, and data feedback loops.
  • Experience reasoning about reward hacking, unstable rewards, distribution shift, stale policies, mode collapse, over-optimization, noisy preferences, and evaluation mismatch.
  • Experience building or operating RL and post-training pipelines at scale with verl, ms-swift, OpenRLHF, equivalent systems, or rollout serving systems such as vLLM.
  • Experience with large-scale training or inference systems, including serving, batching, queueing, GPU utilization, checkpointing, and debugging.
  • Understanding of multimodal post-training for real-time audio-video-language interaction and strong software engineering fundamentals.

Nice to have

  • Experience building post-training systems, RL pipelines, agent training systems, evaluation platforms, or model improvement loops from 0→1.
  • Experience with PPO, GRPO, DPO, online RL, reward modeling, preference data, synthetic data generation, or model-based data improvement.
  • Experience with multimodal, long-context, or real-time interactive systems and mixed training/inference workloads across large GPU clusters.
  • Publications or substantial open-source contributions in RL, post-training, alignment, evaluation, ML systems, or model behavior.

Culture & Benefits

  • Visa sponsorship is available from day one, including O-1, H-1B, and green card support.
  • HSA health plan with approximately $2,000 in annual company contributions.
  • 15 days of PTO, public holidays, and a full week of office closure at year-end.
  • Weekday lunch, drinks, and snacks provided.
  • Commuter benefits and a 401(k) plan.
  • AI-native tooling with unlimited tokens.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →