Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Member of Technical Staff, RL Environments (AI): Building reinforcement learning environments for enterprise AI agents, including realistic tasks, synthetic data, tools, and verifiers with an accent on agent training, evaluation, and capability measurement. Focus on designing reward systems, identifying agent failure points, scaling annotation and synthetic data workflows, and systematically improving environments and model performance.
Location: London; hybrid or remote. The role can also be based from one of the listed offices in Toronto, New York City, San Francisco, Montreal, Paris, Berlin, or Seoul, with no minimum in-office requirement.
Company
Cohere is a security-first enterprise AI company building foundation models and end-to-end AI products for real-world business applications.
What you will do
- Build reinforcement learning environments for different agent capabilities and industry applications.
- Train and evaluate AI agents within those environments.
- Integrate tasks, data, tool implementations, and verifiers into complete evaluation systems.
- Identify gaps in agent performance with modeling and product teams, then improve agents and environments.
- Collaborate with external vendors to create expert-built environments and tools for validating task, data, and verifier quality.
- Automate capability-gap discovery and measure agent performance during training and evaluations.
Requirements
- Experience engineering and optimizing agents for specific industry use cases.
- Experience reviewing agent trajectories, diagnosing failure points, and addressing them through model training or harness engineering.
- Strong focus on measuring agentic capabilities and creating repeatable evaluation processes.
- Experience defining desired agent outcomes, implementing verifiers, and tuning reward designs.
- Experience designing annotation workflows, evaluating agent performance, and verifying data quality.
- Experience building synthetic data pipelines for evaluation and training at scale.
Nice to have
- Experience training with reinforcement learning, including scaling, troubleshooting, and tuning environments.
Culture & Benefits
- Remote-friendly work environment with offices and coworking support for employees who are not near an office.
- Weekly lunch stipend of $75, £75, or the equivalent in local currency.
- Health and dental benefits, mental health budget, and retirement contributions including RRSP matching, 401(k), or pension scheme.
- Six weeks of paid vacation, parental leave top-up for up to six months, and annual enrichment benefits.
- Education and learning stipend, home office stipend, office travel budget for remote employees, and an annual company offsite.
- Office-based employees receive daily lunch, snacks, and community events.
Hiring process
- Applicants may be screened and assessed with AI-enabled tools against the role criteria.
- Accommodations are available during the recruitment process upon request.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →