Назад
Company hidden
11 дней назад

LLM Evaluation Engineer

Формат работы
remote
Тип работы
fulltime
Грейд
middle/senior
Английский
b2
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
LLM Evaluation Engineer (AI): Develop and maintain evaluation infrastructure for large language models, including agent capability evaluations, benchmark design, and failure analysis with an accent on calibration protocols, automated grading, and recurring evaluation workflows. Focus on designing robust benchmarks, analyzing model failures, and shipping reliable tooling that supports research teams from day one.

What you will do

  • Run the full evaluation pipeline end to end and reproduce known results during onboarding.
  • Build and document judge calibration protocols to measure agreement and identify drift zones.
  • Extend existing benchmarks with new tasks targeting capability gaps, including prompt, environment, rubric, automated grader, and QA.
  • Perform failure analysis on model outputs, categorize failure modes, quantify prevalence, and provide recommendations.
  • Own recurring evaluation workflows and ship tooling used by researchers.

Requirements

  • 3+ years experience in software engineering, ML engineering, data science, or research-adjacent roles with concrete evaluation experience.
  • Experience with at least one LLM evaluation framework and hands-on experience with LLMs including prompting and few-shot design.
  • Proficiency in Python, Git, CI/CD, Docker, and Linux command line tools.
  • Understanding of evaluation statistics including Cohen's κ and confidence intervals.
  • Ability to explain LLM-as-judge calibration, perform failure analysis, and knowledge of multiple agent benchmarks.
  • Strong communication skills and comfort with ambiguity.

Nice to have

  • Experience with RLVR / RLHF pipelines.
  • Training data curation and distributed evaluation orchestration experience.
  • Benchmark design from scratch and red teaming/adversarial evaluation experience.
  • Familiarity with psychometrics or measurement theory.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →