Назад
Company hidden
4 часа назад

RL Environment Research Engineer (AI)

Формат работы
hybrid
Тип работы
fulltime
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
RL Environment Research Engineer (AI): Creating, evaluating, and benchmarking reinforcement-learning environments for training AI agents with an accent on agent failure analysis, reward-hacking detection, and scalable evaluation pipelines. Focus on designing benchmark suites, analyzing rollout data, running training experiments, and turning research insights into validated environments and reproducible production workflows.

Location: Hybrid in Mountain View, California, United States

Company

hirify.global is an applied AI research lab focused on curating data and reinforcement-learning environments for training and evaluating AI agents.

What you will do

  • Develop systematic strategies and repeatable recipes for creating high-quality reinforcement-learning environments.
  • Analyze LLM and agent behavior, rollouts, failure modes, reward hacking, and training dynamics.
  • Create benchmark environments for specific agent capabilities and package them for external evaluation platforms.
  • Design and run experiments with small-scale agents to verify environment quality and interpret results.
  • Build and use pipelines for scaling validated environment production and integrating benchmarks into external dashboards.
  • Establish quality standards and evaluation protocols for environments, datasets, and benchmark suites.

Requirements

  • Strong machine-learning foundation through a PhD, MS in ML or computer science, or equivalent industry experience.
  • Proficiency in Python and machine-learning frameworks such as PyTorch or JAX.
  • Experience with reinforcement-learning concepts, agent training, training loops, and experimental design.
  • Ability to analyze complex systems and rollout data, form hypotheses, and identify subtle failure patterns.
  • Experience building pipelines, automation, data-analysis workflows, and reproducible research processes.
  • Comfort working with cloud platforms such as GCP or AWS for experiments at scale.

Nice to have

  • Hands-on experience with reinforcement learning or agent-training systems.
  • Experience with data curation, dataset creation, or evaluation benchmark design.
  • Background in AI safety, robustness testing, or adversarial evaluation.
  • Publications or projects related to reinforcement learning, agent evaluation, or data-centric AI.
  • Experience shipping datasets, benchmarks, or evaluation suites to the community.

Culture & Benefits

  • Health coverage.
  • Flexible work arrangements within the hybrid role structure.
  • Opportunity to shape how the AI community evaluates and trains agents.
  • Research environment supporting diverse research backgrounds and external-facing artifacts.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →