Назад
Company hidden
2 дня назад

Senior AI Evaluation Engineer (LangGraph)

Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Serbia/Ukraine/Poland +7 еще
Вакансия из списка Hirify.GlobalВакансия из Hirify RU Global, списка компаний с восточно-европейскими корнями
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior AI Evaluation Engineer (LangGraph): Building automated evaluation and quality assurance frameworks for agent-based AI systems with an accent on deterministic grading, LLM-as-judge methodologies, and deployment validation. Focus on designing multi-layer testing, CI/CD quality gates, shadow-mode workflows, and production feedback loops that improve agent reliability, accuracy, and safety.

Location: Armenia, Bulgaria, Cyprus, Georgia, Kazakhstan, Latvia, Poland, Romania, Serbia, or Ukraine

Company

hirify.global delivers technology solutions and works on global projects in a collaborative technology environment.

What you will do

  • Design and develop build-time evaluation frameworks and automated test harnesses for LangGraph-based agent systems.
  • Combine deterministic grading and LLM-as-judge methodologies to assess tool selection, execution trajectories, reasoning, outputs, and multi-turn context retention.
  • Build reliability testing using multi-trial, pass-at-k, and pass-power-k approaches.
  • Maintain CI/CD deployment gates, staging validation, shadow-mode comparisons, and controlled rollout evaluation workflows.
  • Integrate production feedback into regression testing and collaborate with engineering and platform teams on evaluation standards, thresholds, and governance.
  • Analyze evaluation results, recommend reliability improvements, and contribute to technical documentation and quality engineering practices.

Requirements

  • 4+ years of experience building automated testing frameworks, evaluation platforms, or quality assurance solutions for ML, LLM, or agent-based systems.
  • Experience designing multi-layer evaluation frameworks with deterministic and LLM-as-judge grading.
  • Experience implementing metric-based deployment quality gates and scalable validation processes for production AI systems.
  • Experience with LangGraph or a comparable agent orchestration framework.
  • Strong Python development skills and knowledge of CI/CD, deployment automation, and release governance.
  • Strong understanding of agent behavior evaluation, AI quality measurement, data analysis, communication, and collaboration.

Nice to have

  • Experience with AWS AgentCore Evaluations and custom evaluator development.
  • Experience with shadow-mode or canary deployment strategies for ML or AI systems.
  • Experience creating production-feedback loops, observability pipelines, and telemetry-driven quality improvements.
  • Knowledge of enterprise AI governance and quality assurance programs.

Culture & Benefits

  • Work on global projects in a supportive, flexible, and innovative technology environment.
  • Vacation, sick pay, and state-holiday time off according to the laws and calendar of the employee’s country.
  • Health insurance support for employees and their loved ones.
  • IT certification cost coverage and access to courses and learning platforms.
  • Corporate events, colleague get-togethers, and technical and everyday workplace support.
  • The benefits package varies by region and contract type.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →