Назад
Company hidden
5 дней назад

AI Evaluation Engineer

263 600 - 395 400$
Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
AI Evaluation Engineer (AI/evaluation infrastructure): Building the execution engines, task-set tooling, grader infrastructure, and reporting systems that make AI model evaluation fast, statistically reliable, and useful for product teams with an accent on offline and online evaluation, human-judge agreement, and production-scale workflows. Focus on designing distributed evaluation systems, measuring confidence and variance, detecting drift, and connecting offline scores with real-world user outcomes.

Location: Remote, Bay Area, California, United States

Salary: $263,600–$395,400 USD annually, depending on location zone, skills, experience, qualifications, and market conditions.

Company

hirify.global builds financial technology products that increase access to the global economy, including Cash App, Square, Afterpay, TIDAL, Bitkey, and Proto.

What you will do

  • Build an execution engine that scores candidate model versions against task sets in minutes.
  • Create tooling to sample production logs, validate tasks, and manage evaluation sets.
  • Develop grader infrastructure using ground-truth checks, rubrics, and LLM-as-judge approaches.
  • Build tools for human-reviewer calibration, judge-to-human agreement measurement, and drift monitoring.
  • Develop leaderboards and reporting with sample sizes, confidence intervals, and run-to-run variance.
  • Connect offline evaluations with online outcomes, in-product feedback, and real conversational signals.

Requirements

  • Experience building production platforms or infrastructure, including distributed batch execution, data pipelines, or systems processing production logs.
  • Strong statistical literacy, including confidence intervals, variance, statistical power, and multiple comparisons.
  • Experience evaluating LLM or ML systems, or deep systems engineering experience with a strong interest in AI evaluation.
  • Ability to build reliable, observable systems and determine whether results represent meaningful improvements or statistical noise.
  • Strong product judgment for internal tools and strong collaboration across product, data, engineering, and ML teams.
  • Must be located in the Bay Area, California, United States.

Culture & Benefits

  • Remote work with a distributed, globally collaborative environment.
  • Medical insurance and flexible time off.
  • Retirement savings plans and modern family-planning benefits.
  • Work focused on improving the speed, accuracy, and reliability of AI product development.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →