Назад
Company hidden
6 часов назад

Inference QA Engineer (AI)

Формат работы
remote (только United_kingdom)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
UK
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Inference QA Engineer (AI) (LLM evaluation and agentic systems): Building evaluation, monitoring, and regression systems for an agentic LLM platform serving energy and heavy-industry operations with an accent on ground-truth scoring, tool routing, and inference reliability. Focus on designing tiered test frameworks, tracing multi-step agent failures, and catching latency, non-termination, hallucination, and run-to-run instability before deployment.

Location: Remote UK; fully remote

Company

hirify.global builds Orbital, a physics-informed foundation model and agentic LLM platform for energy operations, serving oil and gas, refinery, petrochemical, and heavy-industry customers.

What you will do

  • Own and extend tiered evaluation frameworks covering data retrieval, statistical analysis, open-ended inference, and root-cause analysis.
  • Design ground-truth test sets with expected answers, pass criteria, known failure modes, and source tables.
  • Run and maintain large-scale test harnesses against live model endpoints, converting execution data into deployment verdicts.
  • Build monitoring and observability for multi-step agentic systems, including planner routing, tool execution, grounding, latency, and termination.
  • Turn deployment feedback and subject-matter expertise into permanent regression tests.
  • Gate releases by creating failure-mode-focused tests, measuring pass rates, and verifying that defects do not recur.

Requirements

  • 5+ years of experience in software, ML, data, or QA engineering with ownership of a quality-critical system.
  • Strong Python skills and experience working with FastAPI, Postgres, Docker, distributed services, and logs.
  • Fluency with LLM prompting, tool and function calling, agentic loops, RAG, hallucinations, refusals, truncation, and non-determinism.
  • Experience designing evaluations using LLM-as-judge, deterministic checks, ground-truth scoring, and statistical consistency measures.
  • SQL literacy for reviewing generated queries and validating tables and answers.
  • Strong monitoring and observability practice, including dashboards, alerts, and trace inspection.

Nice to have

  • Experience evaluating or red-teaming agentic and multi-tool LLM systems.
  • MLflow or similar trace and experiment tooling.
  • Experience working with technical end users and converting feedback into reproducible tests.
  • Experience with time-series, forecasting, industrial, or operational data.

Culture & Benefits

  • Founding-level ownership of the inference quality function.
  • Mandate to build the assurance layer from the ground up.
  • Direct impact on the reliability of AI systems used by engineers and analysts.
  • Work on a product focused on improving operational efficiency, safety, and carbon intensity in energy.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →