Назад
Company hidden
6 дней назад

AI Engineer (Harness)

Формат работы
hybrid
Тип работы
fulltime
Английский
b2
Страна
UK
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
AI Engineer (Harness) (AI/LLM evaluation): Building experimental evaluation design, benchmark infrastructure, and Python harnesses for AI models and autonomous agentic workloads with an accent on statistical validation, LLM-as-a-judge pipelines, and non-deterministic multi-step system analysis. Focus on calibrating evaluation metrics, analyzing agent failure modes, and improving tool orchestration, retries, context management, and safety guardrails.

Location: London, United Kingdom; hybrid work with attendance at the London office 3 days a week

Company

hirify.global Technologies is a dual-use technology company building secure software, platforms, and infrastructure for government and commercial workflows.

What you will do

  • Design and maintain scalable Python evaluation harnesses for multi-step AI agent systems.
  • Measure task completion, trajectory quality, tool-use correctness, cost, latency, and variance.
  • Apply hypothesis testing, confidence intervals, bootstrapping, and power analysis to validate system improvements.
  • Build and calibrate LLM-as-a-judge pipelines against human annotations and detect evaluation biases.
  • Curate benchmark datasets, annotation guidelines, scoring rubrics, and inter-annotator agreement processes.
  • Analyze agent failures and translate findings into improvements to prompts, tool loops, retries, context management, and safety guardrails.

Requirements

  • Expertise in statistics, experimental design, hypothesis testing, bootstrapping, power analysis, and significance testing for probabilistic machine learning systems.
  • Experience validating LLM judges against human ground truth and tracking agreement metrics.
  • Strong experience designing evaluation datasets, scoring rubrics, label-noise processes, and multi-step system benchmarks.
  • Practical proficiency with LLM APIs, prompt engineering, structured outputs, function and tool calling, retrieval systems, and Python.
  • Understanding of agent planning loops, tool orchestration, memory management, multi-agent coordination, guardrails, retry logic, and state handling.
  • Ability to perform detailed error analysis and write clean, maintainable harness and scaffolding code.

Nice to have

  • Experience with LangChain, LangGraph, AutoGen, lm-evaluation-harness, Promptfoo, or Ragas.
  • Experience running evaluation pipelines or agent workloads in secure enclaves or Kubernetes environments.

Culture & Benefits

  • Hybrid work setup with collaboration in the London office.
  • Equity participation and competitive salary.
  • Employer pension contributions and private health insurance.
  • Company retreats and meetups designed to support work-life balance.
  • Inclusive, compassionate, and flexible working culture.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →