Назад
Company hidden
1 день назад

Senior ML Evaluation Engineer (AWS Agent Evaluation)

Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Serbia/Ukraine/Poland +7 еще
Вакансия из списка Hirify.GlobalВакансия из Hirify RU Global, списка компаний с восточно-европейскими корнями
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior ML Evaluation Engineer (AWS Agent Evaluation): Building evaluation frameworks and scalable quality pipelines for enterprise AI agents and machine learning systems with an accent on LLM evaluation, deterministic Python-based validation, and observability. Focus on integrating CI/CD quality gates, analyzing production evaluation results, and establishing governance standards for reliable AI outcomes.

Location: Armenia, Bulgaria, Cyprus, Georgia, Kazakhstan, Latvia, Poland, Romania, Serbia, or Ukraine

Company

hirify.global delivers software engineering and technology solutions through distributed teams working on global projects.

What you will do

  • Develop and maintain evaluation frameworks for AI agents and machine learning systems.
  • Design LLM-as-a-judge methodologies and custom Python evaluators for deterministic quality and compliance checks.
  • Define enterprise quality standards, assessment dimensions, and pass/fail criteria.
  • Build evaluation workflows for responses, tool invocations, and end-to-end sessions.
  • Integrate OpenTelemetry data and maintain CI/CD quality gates for models and AI agents.
  • Analyze evaluation results, support production monitoring, and recommend improvements to reliability and performance.

Requirements

  • 5+ years of experience in machine learning engineering or AI platform engineering.
  • Hands-on experience designing LLM evaluation frameworks and custom evaluators.
  • Experience building CI/CD deployment gates for machine learning models, AI applications, or agent-based systems.
  • Strong Python development skills and experience with AI quality metrics, automated testing, and evaluation pipelines.
  • Understanding of agent-based architectures and modern AI application development practices.
  • Strong analytical, problem-solving, written, and verbal communication skills.

Nice to have

  • Experience with AWS Agent Evaluation APIs, including evaluation execution and results analysis.
  • Experience integrating AWS Bedrock Guardrails for PII detection and evaluation workflows.
  • Experience with CloudWatch metrics, OpenTelemetry, and observability-based monitoring.
  • Knowledge of enterprise AI governance and compliance programs.
  • Experience evaluating production AI agents and large-scale machine learning systems.

Culture & Benefits

  • Work on global projects in a supportive, flexible, and innovative technology environment.
  • Vacation and state holidays follow the laws and official calendar of the country of employment.
  • Health insurance support for employees and their loved ones.
  • Ten sick days are available without a doctor's note; additional sick leave follows local laws.
  • IT certification costs, courses, and learning platforms are supported.
  • Corporate events, colleague meet-ups, and workplace technical support are provided.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →