Назад
обновлено 8 дней назад

Machine Learning Engineer (Reinforcement Learning)

Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Machine Learning Engineer (Reinforcement Learning): Build and maintain training, evaluation, and deployment loops for a recursive self-improvement ML system with an accent on reward design, feedback signals, and system stability. Focus on designing robust feedback loops, debugging model regressions, and collaborating to ship reliable ML systems.

Location

Location: Menlo Park, CA, On-site

Company

Hippocratic AI is a healthcare-focused AI startup building a safety-first large language model platform to transform patient outcomes globally.

What you will do

  • Build and maintain training, evaluation, and deployment loops emphasizing reproducibility and reliability.
  • Design and implement reward and feedback signals; mitigate reward hacking and specification gaming.
  • Develop evaluation harnesses and metrics to measure model improvements.
  • Own data pipelines and automated data flywheels feeding the learning loop.
  • Debug model-quality regressions and stabilize non-stationary training and feedback loops.
  • Collaborate with research and product teams to deliver robust, shippable ML systems.

Requirements

  • On-site work in Menlo Park, CA
  • Strong machine learning engineering fundamentals with excellent Python and clean, well-tested ML training code.
  • Experience with data pipelines, distributed/large-scale training, and experiment tracking.
  • Hands-on experience with feedback or learning loops such as reward models, evaluation harnesses, or RLHF/RLAIF pipelines.
  • Working knowledge of reinforcement learning foundations including reward modeling and credit assignment.
  • Experience shipping ML systems into production and maintaining system health over time.

Nice to have

  • PhD or MS in RL/ML with production experience.
  • Experience at labs or companies working on RLHF, agents, or large-scale ML infrastructure.
  • Familiarity with LLM fine-tuning, evaluation frameworks, or agent orchestration.

Culture & Benefits

  • Work with a team of AI pioneers, physicians, and healthcare leaders.
  • Backed by leading healthcare and AI investors with significant funding.
  • Equal opportunity employer committed to diversity and inclusion.
  • Support for accommodations during the hiring process.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →