Назад
Company hidden
3 часа назад

Staff Machine Learning Engineer (Agent Eval Platform)

Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Staff Machine Learning Engineer (Agent Eval Platform) (LLM evaluation and reward modeling): Building the judgement layer for an agent evaluation platform that scores multi-step LLM agent trajectories with an accent on rubric design, human calibration, confidence-aware scoring, and deterministic validation. Focus on turning calibrated trajectory judgements into process reward models, optimizing agent behavior in simulation, and diagnosing offline-to-online divergence.

Location: Required in office in Mountain View, California, United States

Company

Moveworks provides an agentic AI assistant platform that connects employees with enterprise systems through natural language, and is part of ServiceNow.

What you will do

  • Build the judgement layer of the agent evaluation platform, including shared judges and per-item evaluation rubrics.
  • Combine deterministic validators for checkable system state with LLM judges for subjective trajectory quality.
  • Calibrate judges against human-labeled trajectories and develop confidence-aware scoring that routes uncertain cases to human review.
  • Fine-tune small judge models and guard against correlated blind spots between simulators and evaluators.
  • Investigate offline-to-online divergence and continuously reseed evaluation suites from production failures.
  • Develop process reward models and use simulation-based signals to optimize prompts, tool selection, planning, retrieval, and routing.

Requirements

  • 8+ years of experience in applied ML, data science, or ML-adjacent engineering, with shipped production work.
  • Experience converting subjective human judgement into reliable measurements that people and models can use.
  • Strong applied ML fundamentals and experience evaluating, prompting, and fine-tuning LLMs.
  • Strong Python skills and the ability to ship production-grade code.
  • Clear communication, strong ownership, comfort with ambiguity, and a bias toward shipping.
  • Experience in at least three areas including LLM-as-judge evaluation, annotation programs, ranking or experimentation evaluation, small-model fine-tuning, reward modeling, RLHF/RLAIF, trajectory analysis, or disciplined prompt engineering.

Culture & Benefits

  • Regular employee position with a full-time schedule.
  • Work persona is assigned according to the role and location, with this position designated as required in office.
  • Work is performed at startup pace with a high degree of ownership and trust.
  • The role addresses an open applied research problem involving LLM-based agents and real enterprise-system actions.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →