Назад
Company hidden
12 дней назад

Senior RL Engineer (Reinforcement Learning)

Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Canada
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior RL Engineer (Reinforcement Learning): Building high-fidelity 2D/3D simulation environments and training autonomous agents with an accent on reward engineering, policy architectures, and multi-sensor observation spaces. Focus on implementing PPO, SAC, and Offline RL algorithms, reducing the sim-to-real gap, and ensuring reliable agent behavior under complex real-world conditions.

Location: Montréal, Québec, Canada

Company

hirify.global is a global media and entertainment company producing and distributing film, television, streaming, news, and themed entertainment content.

What you will do

  • Coordinate with ML engineers, annotation teams, and TPMs to define data, simulation, and training requirements.
  • Build and maintain high-fidelity 2D/3D simulation environments using tools such as Unity, Unreal, or Isaac Sim.
  • Design and tune reward functions that align autonomous-agent behavior with product goals and safety constraints.
  • Develop and optimize reinforcement learning algorithms, including PPO, SAC, and Offline RL, for high-dimensional observation spaces.
  • Analyze the sim-to-real gap and implement domain randomization or adaptation techniques for reliable real-world performance.

Requirements

  • Master’s or PhD in Robotics, Computer Science, AI, or a related field, with a focus on reinforcement learning, imitation learning, or online machine learning.
  • Proven experience as an RL Engineer or Research Engineer in a fast-paced environment.
  • Experience in a multidisciplinary industry such as robotics, smart grids, precision agriculture, game development, or aerospace.
  • Fluency with Python, Git, and the Unix shell.
  • Experience with RL frameworks such as Ray RLlib, Stable Baselines3, or CleanRL, plus physics engines such as MuJoCo or Bullet and 3D game engines.
  • Strong mathematical knowledge of Markov decision processes and gradient-based optimization, with attention to detail when debugging nondeterministic agent behavior.

Culture & Benefits

  • Full-time position within the NBCU Corporate business segment.
  • Work with multidisciplinary ML, annotation, technical program management, and research engineering stakeholders.
  • External candidates may be required to attend an in-person interview at an hirify.global location before a hiring decision.
  • Equal employment opportunities and reasonable accommodation support are provided throughout the application and recruitment process.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →