2 дня назад
Research Engineer (Reinforcement Learning)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Research Engineer (Reinforcement Learning): Build and scale distributed reinforcement learning post-training systems integrating trainers, rollout generation, environments, and reward infrastructure with an accent on managing thousands of GPUs and asynchronous architectures. Focus on designing scalable RL environments, reward verification, and developing tooling for evaluation, monitoring, and debugging large RL runs.
Location
Location: Redwood City, CA (Hybrid)
Company
builds unified general intelligence with a focus on multimodality beyond language models, emphasizing vision and physical world operation.
What you will do
- Design, build, and scale distributed RL post-training systems across thousands of GPUs.
- Develop high-throughput rollout generation integrating inference engines and asynchronous schemes.
- Create RL environments for agentic, multi-step tasks with sandboxed execution and multimodal interaction.
- Build reward infrastructure including verifiable rewards, reward-model serving, and defenses against reward hacking.
- Develop evaluation, monitoring, and debugging tools to maintain stability of large RL runs.
- Advance training efficiency and implement new post-training ideas in production.
Requirements
- Location: Redwood City, CA with hybrid work format.
- Hands-on experience with post-training LLMs using RL at scale (PPO/GRPO-family, RLHF, RLVR).
- Extensive distributed PyTorch training and parallelism for foundation models.
- Experience building RL environments, reward functions, verifiers, and evaluation harnesses for LLM agents.
- Familiarity with RL post-training frameworks and rollout inference engines.
- Strong understanding of GPU clusters, networking, and communication libraries under mixed workloads.
Nice to have
- Experience running RL training across 100+ GPUs with asynchronous or disaggregated architectures.
- Containerization and orchestration skills (Kubernetes, Ray) for large environment fleets.
- Research or open-source contributions in RL for LLMs.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
Cohere
5 дней назад
Member of Technical Staff (AI)
4 дня назад
Senior Research Engineer, LLM Training & Post-Training (PyTorch)
165 000 - 310 000$
6 дней назад
Research Scientist (AI)
6 дней назад
GPU Performance Engineer (AI)
6 дней назад
Research Engineer / Scientist (Robot Learning)
250 000 - 350 000$
6 дней назад