13 дней назад
Member of Technical Staff — RL Research (Experienced)
300 000 - 500 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Member of Technical Staff — RL Research (Experienced) (RL and post-training): Building Nuance’s large-scale RL and post-training stack for real-time audiovisual foundation models with an accent on rollout generation, reward modeling, policy optimization, and multimodal evaluation. Focus on designing distributed training infrastructure, improving interactive behavior and timing, and optimizing throughput, serving latency, GPU utilization, and research iteration speed.
Location: In-person in Seattle, five days a week
Salary: $300,000–$500,000 annual base salary, plus equity
Company
is building photorealistic, real-time AI avatars with emotional intelligence through full-duplex audiovisual foundation models.
What you will do
- Build Nuance’s RL and post-training stack from 0→1 and scale it from 1→10.
- Develop rollout generation, policy optimization, reward and reference model serving, evaluation, checkpointing, observability, and debugging systems.
- Design abstractions connecting research ideas with production-scale trainers, rollout workers, evaluators, data queues, experience buffers, and checkpoint promotion.
- Develop feedback loops for turn-taking, interruption, timing, emotional response, audiovisual coherence, instruction following, and real-time interaction quality.
- Optimize rollout throughput, serving latency, GPU utilization, policy updates, queueing, checkpoint overhead, and research iteration speed.
Requirements
- Significant hands-on experience with RL, RLHF, RLAIF, post-training, alignment, or large-scale fine-tuning for foundation models.
- Deep understanding of policy optimization, reward modeling, preference optimization, rejection sampling, KL control, evaluation, and data feedback loops.
- Experience reasoning about reward hacking, unstable rewards, distribution shift, stale policies, mode collapse, over-optimization, noisy preferences, and evaluation mismatch.
- Experience building or operating RL and post-training pipelines at scale with verl, ms-swift, OpenRLHF, equivalent systems, or rollout serving systems such as vLLM.
- Experience with large-scale training or inference systems, including serving, batching, queueing, GPU utilization, checkpointing, and debugging.
- Understanding of multimodal post-training for real-time audio-video-language interaction and strong software engineering fundamentals.
Nice to have
- Experience building post-training systems, RL pipelines, agent training systems, evaluation platforms, or model improvement loops from 0→1.
- Experience with PPO, GRPO, DPO, online RL, reward modeling, preference data, synthetic data generation, or model-based data improvement.
- Experience with multimodal, long-context, or real-time interactive systems and mixed training/inference workloads across large GPU clusters.
- Publications or substantial open-source contributions in RL, post-training, alignment, evaluation, ML systems, or model behavior.
Culture & Benefits
- Visa sponsorship is available from day one, including O-1, H-1B, and green card support.
- HSA health plan with approximately $2,000 in annual company contributions.
- 15 days of PTO, public holidays, and a full week of office closure at year-end.
- Weekday lunch, drinks, and snacks provided.
- Commuter benefits and a 401(k) plan.
- AI-native tooling with unlimited tokens.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
Snowflake
9 дней назад
AI Research Scientist (New Grad)
176 000 - 230 000$
11 дней назад
AI Researcher, LLMs
200 000 - 300 000$
Scale AI
11 дней назад
Staff Software Engineer, RL Environments
252 000 - 315 000$
Waymo
10 дней назад
Staff Machine Learning Engineer – VLM/LLM Staff Research Scientist, Perception Machine Learning Engineer, Prediction & Planning Foundation Model Data, Software Engineer (AI)
238 000 - 302 000$
10 дней назад
Staff VLA Engineer (Autonomous Driving AI)
189 000 - 311 220$
12 дней назад
Senior/Staff Machine Learning Engineer (Reinforcement Learning), Motion Planning
130 000 - 220 000$