12 часов назад
Member of Technical Staff, Post-Training (AI)
180 000 - 450 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Member of Technical Staff, Post-Training (AI): Developing post-training strategies and agentic reinforcement learning systems for coding, computer use, and long-horizon task completion with an accent on simulation environments, reward modeling, and synthetic data generation. Focus on designing rigorous experiments, scaling model training, and building evaluations for code correctness, execution success, and multi-step tool use.
Location: San Jose, United States
Salary: $180,000–$450,000 annually
Company
is an artificial intelligence company developing personalized, proactive, multimodal intelligence and next-generation hardware as a unified interface between humans and machines.
What you will do
- Design and implement primarily reinforcement-learning-based post-training strategies for coding agents with multi-step reasoning, tool use, and long-horizon task completion.
- Build and scale code execution sandboxes, computer-use environments, tool-calling harnesses, and verifiable reward systems for agentic reinforcement learning.
- Develop outcome-based, execution-based, and process-based reward modeling pipelines.
- Scale synthetic data generation and trajectory distillation pipelines to improve reinforcement learning sample efficiency.
- Run ablation studies and diagnose interactions between algorithms, data mixtures, reward shaping, and scale.
- Build evaluations for code correctness, execution success, and multi-step tool use while collaborating with mid-training, infrastructure, and product teams.
Requirements
- Strong machine learning background with hands-on experience training or fine-tuning large language, multimodal, or equivalent models.
- Deep understanding of reinforcement learning, including policy optimization, reward design, exploration, and environment design.
- Experience with simulation or execution environments such as code interpreters, sandboxes, game environments, or robotics simulators.
- Ability to design rigorous experiments, diagnose training failures, and identify scaling bottlenecks.
- Proficiency in Python and PyTorch, with comfort working across research and systems code.
- Ability to work in a fast-moving, research-focused environment where approaches may be uncertain initially.
Nice to have
- Experience applying RL algorithms such as RLHF, DPO, GRPO, or PPO to language or code.
- Familiarity with coding-agent benchmarks including SWE-bench, HumanEval, or LiveCodeBench.
- Experience with reward modeling, trajectory-based training, imitation learning, or data distillation.
- Experience with computer-use agents, GUI agents, or tool-using language models.
- Experience training or scaling models with 10B+ parameters, or contributions to open-source ML projects and research publications.
Culture & Benefits
- Research-forward environment focused on emerging agentic AI capabilities.
- Opportunity to work at the intersection of reinforcement learning, simulation, and large-scale model training.
- Full-time employment with an annual US base salary range of $180,000–$450,000.
- Total compensation may include additional components and benefits depending on the role.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
6 часов назад
Field Team - Member of Technical Staff (AI)
200 000 - 325 000$
9 часов назад
AI Researcher
160 000 - 300 000$
7 часов назад
Member of Technical Staff (AI)
4 часа назад
Member of Technical Staff (Research)
6 часов назад
Software Engineer (AI Inference & RL Systems)
225 000 - 550 000$
4 часа назад