Назад
Company hidden
5 дней назад

Evaluation Engineer (AI)

Формат работы
remote (только USA)/hybrid/onsite
Тип работы
fulltime
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Evaluation Engineer (AI): Building and maintaining trusted, versioned evaluation datasets for AI and LLM pipeline components with an accent on defining correctness, sourcing reliable labels, and validating automated graders. Focus on measuring model quality, detecting contamination and leakage, calibrating LLM-as-judge systems, and translating system logic into actionable data for engineering teams.

Location: US (Remote); candidates must be authorized to work in the United States without current or future employment visa sponsorship.

Company

An AI-native real estate technology company building a platform that helps homebuilders, developers, and investors find, analyze, and act on land opportunities using proprietary technologies across all 50 states.

What you will do

  • Decompose AI and LLM pipelines into evaluable modules with explicit input-to-output contracts.
  • Define correctness through rubrics, label schemas, edge-case policies, and business-relevant tolerances.
  • Build trusted evaluation datasets using production samples, human labeling, synthetic or oracle-generated data, programmatic cases, and adversarial examples.
  • Version datasets, track lineage and model or prompt exposure, prevent contamination and leakage, and refresh stale examples.
  • Calibrate automated graders against human gold sets and report precision, recall, F1, confusion matrices, calibration, confidence intervals, and sample-size requirements.
  • Partner with engineering and ML teams to integrate datasets into evaluation harnesses and identify data-related model failures.

Requirements

  • Production experience building or evaluating ML or LLM-powered systems where output quality was a responsibility.
  • Previous experience creating evaluation or validation datasets and explaining how correctness and label quality were established.
  • Fluency in ML validation fundamentals, including train/validation/test discipline, stratified sampling, class imbalance, calibration, inter-rater agreement, significance testing, and statistical power.
  • Strong Python and SQL skills, with the ability to pull and reshape data independently.
  • Hands-on experience with prompting, structured outputs, agent and tool-use harnesses, and LLM failure modes such as nondeterminism, prompt sensitivity, and evaluator bias.
  • Must be authorized to work in the United States without current or future sponsorship.

Nice to have

  • End-to-end experience running human-labeling programs, including vendor selection, guidelines, QA sampling, and cost-quality trade-offs.
  • Experience with evaluation tooling or data-labeling platforms.
  • Experience using frontier models as distillation or oracle sources and evaluating where that approach breaks down.

Culture & Benefits

  • Full-time employment with medical, dental, and vision coverage for employees and partial dependent coverage.
  • Competitive salary, early-stage equity, unlimited PTO, and a hybrid or remote stipend.
  • Remote, hybrid, or in-office flexibility, with an office headquarters in Portland, Oregon.
  • Budget for labeling vendors, oracle compute, and intra-office travel.
  • Two to three annual in-person team meetups.
  • Culture centered on truth, precision, velocity, and ownership.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →