Назад
Company hidden
10 часов назад

Manager; AI Evaluation Engineering

147 760 - 240 110$
Формат работы
onsite
Тип работы
fulltime
Грейд
lead
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Manager; AI Evaluation Engineering (GenAI): Leading evaluation and validation of generative AI products including intelligent agents, digital assistants, and other AI-driven experiences with an accent on software quality, test automation, and reliability. Focus on designing evaluation methodologies for non-deterministic systems, building observability and monitoring practices, and scaling AI engineering standards across enterprise products.

Location: Broomfield, Colorado; Chicago or Peoria, Illinois; or Irving, Texas, United States

Salary: $147,760–$240,110 per year

Company

hirify.global develops products and services for construction, infrastructure, and other industrial markets, with a focus on technology, digital solutions, and data.

What you will do

  • Lead a team responsible for evaluating and validating generative AI solutions, including intelligent agents and digital assistants.
  • Provide technical direction and support while aligning evaluation work with company goals.
  • Manage individual and team performance, identifying training and development needs.
  • Own the quality, robustness, and reliability of AI engineering products.
  • Establish and supervise engineering best practices across development processes.

Requirements

  • Experience leading software quality, test automation, validation, and AI evaluation teams.
  • Experience scaling processes, tools, metrics, and engineering practices across multiple products and enterprise initiatives.
  • Experience delivering enterprise-scale software and GenAI solutions in hybrid cloud and embedded or edge environments.
  • Expertise in GenAI architecture, LLMs, SLMs, multimodal and real-time models, prompt engineering, agentic systems, RAG, vector databases, embeddings, chunking, and fine-tuning.
  • Knowledge of AI evaluation methods, including metric design, RAG quality assessment, safety evaluation, A/B testing, human-in-the-loop validation, and testing of non-deterministic systems.
  • Knowledge of software development, the software development life cycle, quality assurance, and testing practices.

Nice to have

  • Experience with Azure, AWS, GCP, Azure AI Foundry, SageMaker, Bedrock, Snowflake Cortex, LangChain, LangGraph, Langfuse, Arize, LangSmith, Humanloop, Ragas, DeepEval, Phoenix, or similar technologies.
  • Experience with CI/CD, automated testing, incident management, feature flags, blue/green deployments, and canary deployments.
  • Experience with AI observability, tracing, telemetry, root-cause analysis, synthetic data generation, automated test generation, and LLM-assisted evaluation.
  • Familiarity with MCP and A2A standards.

Culture & Benefits

  • Medical, dental, and vision benefits.
  • Paid time off, holidays, and volunteer time.
  • 401(k) savings plan, HSA, and FSA.
  • Career development, tuition reimbursement, and incentive bonus opportunities.
  • Disability, life insurance, parental leave, adoption benefits, employee assistance, and employee discounts.
  • Visa sponsorship is not available for this position.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →