2 дня назад
Senior AI Evaluation Engineer (LangGraph)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior AI Evaluation Engineer (LangGraph): Building automated evaluation and quality assurance frameworks for agent-based AI systems with an accent on deterministic grading, LLM-as-judge methodologies, and deployment validation. Focus on designing multi-layer testing, CI/CD quality gates, shadow-mode workflows, and production feedback loops that improve agent reliability, accuracy, and safety.
Location: Armenia, Bulgaria, Cyprus, Georgia, Kazakhstan, Latvia, Poland, Romania, Serbia, or Ukraine
Company
delivers technology solutions and works on global projects in a collaborative technology environment.
What you will do
- Design and develop build-time evaluation frameworks and automated test harnesses for LangGraph-based agent systems.
- Combine deterministic grading and LLM-as-judge methodologies to assess tool selection, execution trajectories, reasoning, outputs, and multi-turn context retention.
- Build reliability testing using multi-trial, pass-at-k, and pass-power-k approaches.
- Maintain CI/CD deployment gates, staging validation, shadow-mode comparisons, and controlled rollout evaluation workflows.
- Integrate production feedback into regression testing and collaborate with engineering and platform teams on evaluation standards, thresholds, and governance.
- Analyze evaluation results, recommend reliability improvements, and contribute to technical documentation and quality engineering practices.
Requirements
- 4+ years of experience building automated testing frameworks, evaluation platforms, or quality assurance solutions for ML, LLM, or agent-based systems.
- Experience designing multi-layer evaluation frameworks with deterministic and LLM-as-judge grading.
- Experience implementing metric-based deployment quality gates and scalable validation processes for production AI systems.
- Experience with LangGraph or a comparable agent orchestration framework.
- Strong Python development skills and knowledge of CI/CD, deployment automation, and release governance.
- Strong understanding of agent behavior evaluation, AI quality measurement, data analysis, communication, and collaboration.
Nice to have
- Experience with AWS AgentCore Evaluations and custom evaluator development.
- Experience with shadow-mode or canary deployment strategies for ML or AI systems.
- Experience creating production-feedback loops, observability pipelines, and telemetry-driven quality improvements.
- Knowledge of enterprise AI governance and quality assurance programs.
Culture & Benefits
- Work on global projects in a supportive, flexible, and innovative technology environment.
- Vacation, sick pay, and state-holiday time off according to the laws and calendar of the employee’s country.
- Health insurance support for employees and their loved ones.
- IT certification cost coverage and access to courses and learning platforms.
- Corporate events, colleague get-togethers, and technical and everyday workplace support.
- The benefits package varies by region and contract type.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →