6 дней назад
AI Engineer (Harness)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
AI Engineer (Harness) (AI/LLM evaluation): Building experimental evaluation design, benchmark infrastructure, and Python harnesses for AI models and autonomous agentic workloads with an accent on statistical validation, LLM-as-a-judge pipelines, and non-deterministic multi-step system analysis. Focus on calibrating evaluation metrics, analyzing agent failure modes, and improving tool orchestration, retries, context management, and safety guardrails.
Location: London, United Kingdom; hybrid work with attendance at the London office 3 days a week
Company
Technologies is a dual-use technology company building secure software, platforms, and infrastructure for government and commercial workflows.
What you will do
- Design and maintain scalable Python evaluation harnesses for multi-step AI agent systems.
- Measure task completion, trajectory quality, tool-use correctness, cost, latency, and variance.
- Apply hypothesis testing, confidence intervals, bootstrapping, and power analysis to validate system improvements.
- Build and calibrate LLM-as-a-judge pipelines against human annotations and detect evaluation biases.
- Curate benchmark datasets, annotation guidelines, scoring rubrics, and inter-annotator agreement processes.
- Analyze agent failures and translate findings into improvements to prompts, tool loops, retries, context management, and safety guardrails.
Requirements
- Expertise in statistics, experimental design, hypothesis testing, bootstrapping, power analysis, and significance testing for probabilistic machine learning systems.
- Experience validating LLM judges against human ground truth and tracking agreement metrics.
- Strong experience designing evaluation datasets, scoring rubrics, label-noise processes, and multi-step system benchmarks.
- Practical proficiency with LLM APIs, prompt engineering, structured outputs, function and tool calling, retrieval systems, and Python.
- Understanding of agent planning loops, tool orchestration, memory management, multi-agent coordination, guardrails, retry logic, and state handling.
- Ability to perform detailed error analysis and write clean, maintainable harness and scaffolding code.
Nice to have
- Experience with LangChain, LangGraph, AutoGen, lm-evaluation-harness, Promptfoo, or Ragas.
- Experience running evaluation pipelines or agent workloads in secure enclaves or Kubernetes environments.
Culture & Benefits
- Hybrid work setup with collaboration in the London office.
- Equity participation and competitive salary.
- Employer pension contributions and private health insurance.
- Company retreats and meetups designed to support work-life balance.
- Inclusive, compassionate, and flexible working culture.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →