4 дня назад
Senior AI Evals Engineer (AI Quality & Safety) - Libra-Legal AI Assistant (m/w/d) (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior AI Evals Engineer (AI Quality & Safety) (AI): Building an evaluation platform for a legal AI assistant with an accent on datasets, LLM-as-judge rubrics, regression gates, trace analysis, and automated optimization. Focus on defining defensible quality measures across legal tasks and jurisdictions, protecting privileged data, and designing guardrails against hallucinated citations, leakage, and prompt injection.
Location: Berlin, Germany; hybrid work with in-person interviews and possible onsite attendance at a office.
Company
Libra is an independent AI-focused business unit within 's Legal & Regulatory division, developing generative AI solutions for legal professionals.
What you will do
- Build and own a self-service evaluation platform using Python, FastAPI, and Langfuse, including datasets, judges, harnesses, regression gates, and cost-quality monitoring.
- Define quality measures for legal research, drafting, summarisation, and retrieval across different jurisdictions.
- Design version-controlled rubrics and recalibrate LLM-as-judge systems with Legal Engineers and subject-matter experts.
- Analyze results and traces to identify failures, improvement levers, and automated fixes.
- Run automated prompt and pipeline optimization with GEPA, DSPy, and trusted fitness functions.
- Design and red-team guardrails against ungrounded advice, fabricated or misattributed citations, jurisdiction and language leakage, and prompt injection.
Requirements
- Bachelor's degree or equivalent in Computer Science, Software Engineering, Statistics, Data Science, or a related technical field.
- At least 5 years of software engineering experience, including at least 1 year building production LLM-powered products.
- Strong Python, FastAPI, pandas, NumPy, and notebook-based analysis skills.
- Hands-on experience designing evaluations, including datasets, rubrics, LLM-as-judge systems, benchmarking, and human labelling.
- Security and data-privacy experience for traces, documents, and datasets derived from privileged legal material.
- Excellent English communication skills; genuine interest in the legal domain.
Nice to have
- Experience with automated prompt or pipeline optimization using GEPA, DSPy, or similar tools.
- Broader data science experience, such as embedding clustering for dataset coverage analysis.
- Advanced technical degree.
Culture & Benefits
- Lean, agile environment with close collaboration across AI Engineering, Legal Engineering, and Product.
- Startup-like speed and ownership backed by the scale and stability of an established global organization.
- Work based at the Merantix AI Campus in Berlin, alongside a community of AI innovators.
- AI coding agents such as Claude Code, Codex, and Cursor are part of the daily workflow.
Hiring process
- Interviews must be completed without AI tools, external prompts, or third-party assistance.
- In-person interviews are included; applicants may be required to appear onsite at a office.
- Use of AI-generated responses or third-party support during interviews may result in disqualification.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →