5 дней назад
Evaluation Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Evaluation Engineer (AI): Building and maintaining trusted, versioned evaluation datasets for AI and LLM pipeline components with an accent on defining correctness, sourcing reliable labels, and validating automated graders. Focus on measuring model quality, detecting contamination and leakage, calibrating LLM-as-judge systems, and translating system logic into actionable data for engineering teams.
Location: US (Remote); candidates must be authorized to work in the United States without current or future employment visa sponsorship.
Company
An AI-native real estate technology company building a platform that helps homebuilders, developers, and investors find, analyze, and act on land opportunities using proprietary technologies across all 50 states.
What you will do
- Decompose AI and LLM pipelines into evaluable modules with explicit input-to-output contracts.
- Define correctness through rubrics, label schemas, edge-case policies, and business-relevant tolerances.
- Build trusted evaluation datasets using production samples, human labeling, synthetic or oracle-generated data, programmatic cases, and adversarial examples.
- Version datasets, track lineage and model or prompt exposure, prevent contamination and leakage, and refresh stale examples.
- Calibrate automated graders against human gold sets and report precision, recall, F1, confusion matrices, calibration, confidence intervals, and sample-size requirements.
- Partner with engineering and ML teams to integrate datasets into evaluation harnesses and identify data-related model failures.
Requirements
- Production experience building or evaluating ML or LLM-powered systems where output quality was a responsibility.
- Previous experience creating evaluation or validation datasets and explaining how correctness and label quality were established.
- Fluency in ML validation fundamentals, including train/validation/test discipline, stratified sampling, class imbalance, calibration, inter-rater agreement, significance testing, and statistical power.
- Strong Python and SQL skills, with the ability to pull and reshape data independently.
- Hands-on experience with prompting, structured outputs, agent and tool-use harnesses, and LLM failure modes such as nondeterminism, prompt sensitivity, and evaluator bias.
- Must be authorized to work in the United States without current or future sponsorship.
Nice to have
- End-to-end experience running human-labeling programs, including vendor selection, guidelines, QA sampling, and cost-quality trade-offs.
- Experience with evaluation tooling or data-labeling platforms.
- Experience using frontier models as distillation or oracle sources and evaluating where that approach breaks down.
Culture & Benefits
- Full-time employment with medical, dental, and vision coverage for employees and partial dependent coverage.
- Competitive salary, early-stage equity, unlimited PTO, and a hybrid or remote stipend.
- Remote, hybrid, or in-office flexibility, with an office headquarters in Portland, Oregon.
- Budget for labeling vendors, oracle compute, and intra-office travel.
- Two to three annual in-person team meetups.
- Culture centered on truth, precision, velocity, and ownership.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
9 дней назад
Agent Architect (AI)
125 000 - 220 000$
8 дней назад
Senior Applied AI Engineer (LLM)
200 000 - 240 000$
6 дней назад
Software Engineer (AI)
110 000 - 138 000$
9 дней назад
Staff Engineer (Applied AI / ML)
Anthropic
5 дней назад
Applied AI Research Engineer
300 000 - 400 000$
8 дней назад