6 часов назад
Inference QA Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Inference QA Engineer (AI) (LLM evaluation and agentic systems): Building evaluation, monitoring, and regression systems for an agentic LLM platform serving energy and heavy-industry operations with an accent on ground-truth scoring, tool routing, and inference reliability. Focus on designing tiered test frameworks, tracing multi-step agent failures, and catching latency, non-termination, hallucination, and run-to-run instability before deployment.
Location: Remote UK; fully remote
Company
builds Orbital, a physics-informed foundation model and agentic LLM platform for energy operations, serving oil and gas, refinery, petrochemical, and heavy-industry customers.
What you will do
- Own and extend tiered evaluation frameworks covering data retrieval, statistical analysis, open-ended inference, and root-cause analysis.
- Design ground-truth test sets with expected answers, pass criteria, known failure modes, and source tables.
- Run and maintain large-scale test harnesses against live model endpoints, converting execution data into deployment verdicts.
- Build monitoring and observability for multi-step agentic systems, including planner routing, tool execution, grounding, latency, and termination.
- Turn deployment feedback and subject-matter expertise into permanent regression tests.
- Gate releases by creating failure-mode-focused tests, measuring pass rates, and verifying that defects do not recur.
Requirements
- 5+ years of experience in software, ML, data, or QA engineering with ownership of a quality-critical system.
- Strong Python skills and experience working with FastAPI, Postgres, Docker, distributed services, and logs.
- Fluency with LLM prompting, tool and function calling, agentic loops, RAG, hallucinations, refusals, truncation, and non-determinism.
- Experience designing evaluations using LLM-as-judge, deterministic checks, ground-truth scoring, and statistical consistency measures.
- SQL literacy for reviewing generated queries and validating tables and answers.
- Strong monitoring and observability practice, including dashboards, alerts, and trace inspection.
Nice to have
- Experience evaluating or red-teaming agentic and multi-tool LLM systems.
- MLflow or similar trace and experiment tooling.
- Experience working with technical end users and converting feedback into reproducible tests.
- Experience with time-series, forecasting, industrial, or operational data.
Culture & Benefits
- Founding-level ownership of the inference quality function.
- Mandate to build the assurance layer from the ground up.
- Direct impact on the reliability of AI systems used by engineers and analysts.
- Work on a product focused on improving operational efficiency, safety, and carbon intensity in energy.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →