4 дня назад
RL Environments Engineer (AI)
250 000 - 300 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
RL Environments Engineer (AI): Building high-fidelity coding environments, task graders, and sandboxed execution systems for training and evaluating frontier coding agents with an accent on real codebases, deterministic grading, and adversarial robustness. Focus on designing rigorous tasks, analyzing model failures, automating environment production, and closing loopholes that allow agents to pass without doing the work.
Location: Mountain View, CA — onsite
Base salary: $250,000–$300,000 USD per year; 25% performance-based bonus and equity.
Company
is an applied AI research lab focused on curating data and reinforcement learning environments for training and evaluating agents.
What you will do
- Build high-fidelity coding environments around real codebases, including their conventions, dependencies, tooling, and historical complexity.
- Design tasks across the full lifecycle: prompts, environments, graders, frontier-model execution, failure analysis, and iteration.
- Create deterministic, sandboxed grading and code-execution systems that are resistant to gaming.
- Run frontier coding agents against environments, evaluate their output, and identify silent failures.
- Develop internal tooling and automation that increases the team's environment-production capacity.
- Choose valuable environments targeting areas where frontier models measurably struggle.
Requirements
- Demonstrated experience shipping agentic coding tasks or environments at meaningful volume, with measurable output and production cost.
- Strong software engineering fundamentals and fluency in several programming languages.
- Production experience with large codebases, build systems, testing, deployment, on-call, and root cause analysis.
- An adversarial approach to graders and a clear understanding of how frontier coding agents behave and cut corners.
- Ability to build, debug, and ship independently with limited supervision.
- Availability to work onsite in Mountain View, California.
Nice to have
- Experience with RL training systems, post-training, verifiers, or tool-use harnesses.
- Background in developer tooling, CI/CD, sandboxes, or code-execution infrastructure.
- Experience with automated test generation, fuzzing harnesses, or benchmark suites.
- Contributions to public agentic benchmarks such as Terminal-Bench.
- Open-source work used by other people.
Culture & Benefits
- Health, dental, and vision coverage.
- 401(k) plan.
- Daily onsite lunch.
- Visa sponsorship and relocation support available.
- Direct impact on how the industry trains and evaluates agents.
- Additional compensation includes a performance-based bonus and equity.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
3 дня назад
Research Engineer (QC Automation)
150 000 - 250 000$
2 дня назад
Senior AI Engineer (AI)
150 000 - 220 000$
4 дня назад
Staff Engineer AI/ML (AI)
151 000 - 178 500$
5 дней назад
Software Engineer (AI)
167 000 - 184 000$
2 дня назад
Staff Engineer — Agentic AI
160 000 - 250 000$
3 дня назад
Forward Deployed Engineer (AI/ML)
150 000 - 200 000$