1 день назад
Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Engineer (AI) (LLM evaluation and observability): Building production-grade evaluation, data infrastructure, and observability systems for multimodal retrieval agents used in complex audit workflows with an accent on quality metrics, reliability, and cost optimization. Focus on analyzing agent traces, creating automated quality gates, and improving distributed retrieval and reasoning systems.
Location: On-site in Berlin, Germany
Company
An early-stage AI startup building software and specialized AI agents to automate repetitive work in document-heavy audits.
What you will do
- Build online and offline evaluation systems for LLM agents using golden datasets, ground-truth data, human review workflows, and experiment results.
- Create automated quality gates for changes to prompts, context, models, and agent logic before production release.
- Analyze large volumes of agent traces in analytical and columnar databases to identify failure modes, quality regressions, latency issues, reliability gaps, and cost optimization opportunities.
- Build data retention and replay mechanisms for long-term analysis of production agent behavior.
- Manage observability tools for tracing, monitoring, debugging, and experiment management.
- Collaborate with backend engineers to improve the speed and reliability of retrieval and reasoning agents.
Requirements
- Strong Python and/or backend engineering experience.
- Understanding of LLM and agent evaluation methods, including deterministic checks, ground truth, LLM-as-judge, human review, and quality metrics.
- Experience deploying and operating cloud systems, ideally on GCP.
- Hands-on experience with retrieval or ML pipeline evaluation systems and LLM observability or experimentation tools such as Braintrust, MLflow, Langfuse, or Weights & Biases.
- Comfort working with analytical databases, data warehouses, columnar stores, and high-volume event or trace data.
- Senior-level engineering judgment, including system design, architectural decisions, reliability, observability, monitoring, logging, debugging, and operational trade-offs.
Nice to have
- Experience designing ETL/ELT workflows, event-processing systems, or feedback loops for production data.
- Experience building infrastructure for LLM products or agentic systems, including LLM usage, context window, reasoning token, or model selection optimization.
- Experience with production traces from complex distributed systems or internal engineering platforms.
- Experience with Temporal or similar workflow orchestration systems.
- Background in audit, finance, compliance, early-stage startups, or fast-moving engineering environments.
Culture & Benefits
- Opportunity to shape strategy and build infrastructure at a scaling AI startup.
- Culture centered on first-principles thinking, speed, trust, kindness, and bold ideas.
- Competitive salary with significant equity.
- Generous coding tools budget, flexible vacation, team lunches, and retreats.
- Work from a central Berlin office with close collaboration.
Hiring process
- Introductory call.
- Technical interview followed by a culture deep dive with a co-founder.
- On-site day in Berlin with the team and work on a real problem.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →