4 часа назад
RL Environment Research Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
RL Environment Research Engineer (AI): Creating, evaluating, and benchmarking reinforcement-learning environments for training AI agents with an accent on agent failure analysis, reward-hacking detection, and scalable evaluation pipelines. Focus on designing benchmark suites, analyzing rollout data, running training experiments, and turning research insights into validated environments and reproducible production workflows.
Location: Hybrid in Mountain View, California, United States
Company
is an applied AI research lab focused on curating data and reinforcement-learning environments for training and evaluating AI agents.
What you will do
- Develop systematic strategies and repeatable recipes for creating high-quality reinforcement-learning environments.
- Analyze LLM and agent behavior, rollouts, failure modes, reward hacking, and training dynamics.
- Create benchmark environments for specific agent capabilities and package them for external evaluation platforms.
- Design and run experiments with small-scale agents to verify environment quality and interpret results.
- Build and use pipelines for scaling validated environment production and integrating benchmarks into external dashboards.
- Establish quality standards and evaluation protocols for environments, datasets, and benchmark suites.
Requirements
- Strong machine-learning foundation through a PhD, MS in ML or computer science, or equivalent industry experience.
- Proficiency in Python and machine-learning frameworks such as PyTorch or JAX.
- Experience with reinforcement-learning concepts, agent training, training loops, and experimental design.
- Ability to analyze complex systems and rollout data, form hypotheses, and identify subtle failure patterns.
- Experience building pipelines, automation, data-analysis workflows, and reproducible research processes.
- Comfort working with cloud platforms such as GCP or AWS for experiments at scale.
Nice to have
- Hands-on experience with reinforcement learning or agent-training systems.
- Experience with data curation, dataset creation, or evaluation benchmark design.
- Background in AI safety, robustness testing, or adversarial evaluation.
- Publications or projects related to reinforcement learning, agent evaluation, or data-centric AI.
- Experience shipping datasets, benchmarks, or evaluation suites to the community.
Culture & Benefits
- Health coverage.
- Flexible work arrangements within the hybrid role structure.
- Opportunity to shape how the AI community evaluates and trains agents.
- Research environment supporting diverse research backgrounds and external-facing artifacts.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →