Назад
Company hidden
13 дней назад

Member of Technical Staff — RL Research (New PhD Grad)

250 000 - 350 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
junior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Member of Technical Staff — RL Research (New PhD Grad) (RL/post-training for omni models): Building Nuance’s reinforcement learning and post-training stack for real-time audiovisual AI avatars with an accent on rollout generation, reward modeling, policy optimization, and evaluation. Focus on designing distributed training systems, improving multimodal interactive behavior, and optimizing latency, throughput, GPU utilization, and research iteration speed.

Location: In-person in Seattle, five days a week

Salary: $250,000–$350,000 base salary per year, plus meaningful equity

Company

hirify.global is a research company building photorealistic, real-time AI avatars with emotional intelligence and full-duplex audiovisual interaction.

What you will do

  • Build the RL and post-training stack from 0→1 and scale it from 1→10, including rollout generation, policy optimization, reward and reference model serving, evaluation, checkpointing, and observability.
  • Develop post-training methods including PPO, GRPO, DPO, rejection sampling, RLHF/RLAIF, online RL, and model-based data improvement.
  • Design systems abstractions connecting research ideas to production-scale training runs, including trainers, rollout workers, evaluators, data queues, experience buffers, and checkpoint promotion.
  • Build feedback and evaluation loops for turn-taking, interruption, timing, emotional response, audiovisual coherence, instruction following, and real-time interaction quality.
  • Optimize rollout throughput, serving latency, GPU utilization, policy updates, queueing, checkpointing, and research iteration speed.
  • Evolve the platform as algorithms, model architectures, reward definitions, data sources, and evaluation methods change.

Requirements

  • Completed or near-completion PhD in ML, RL, or a related field, demonstrated through publications, a strong lab or advisor, or substantial open-source work.
  • Strong understanding of policy optimization, reward modeling, preference optimization, rejection sampling, KL control, evaluation, and data feedback loops.
  • Ability to analyze reward hacking, unstable rewards, distribution shift, stale policies, mode collapse, over-optimization, noisy preferences, and evaluation mismatch.
  • Exposure to RL or post-training pipelines through research, internships, or open source, with frameworks such as verl, ms-swift, OpenRLHF, or equivalent, and familiarity with rollout serving systems such as vLLM.
  • Strong software engineering fundamentals and willingness to build reliable systems rather than prototypes only.
  • Visa sponsorship is available from day one, including O-1, H-1B, and green card support.

Nice to have

  • Experience with omni or multimodal post-training for audio-video-language models, especially long-context or real-time interactive systems.
  • Experience with PPO, GRPO, DPO, online RL, RLHF/RLAIF, reward modeling, preference data, synthetic data generation, or model-based data improvement.
  • Prior 0→1 experience building post-training systems, RL pipelines, agent training systems, evaluation platforms, or model improvement loops.
  • Experience with distributed pretraining, data infrastructure, inference serving, simulation, feedback collection, or evaluation infrastructure.
  • Publications or substantial open-source contributions in RL, post-training, alignment, evaluation, ML systems, or model behavior.

Culture & Benefits

  • Research-focused environment working on unsolved problems in real-time, full-duplex AI.
  • Meaningful equity designed for long-term ownership.
  • HSA health plan with approximately $2,000 in annual company contributions.
  • 15 days of PTO, public holidays, and a full week of office closure at year-end.
  • Weekday lunch, drinks, snacks, commuter benefits, and a 401(k) plan.
  • AI-native tooling with unlimited tokens.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →