Назад
Company hidden
7 дней назад

AI Evaluation Engineers

150 000 - 250 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
middle
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
AI Evaluation Engineers (Python/LLM Evaluation): Building production evaluation frameworks, test suites, and offline and online pipelines for customer-facing AI systems with an accent on measurable quality, human-aligned grading, and reliable iteration. Focus on designing golden test cases, calibrating LLM-based graders, translating expert judgment into code, and connecting evaluation results to prompts, agents, model selection, and release readiness.

Location: Hybrid in San Francisco or New York, with 3+ days per week in-office from Tuesday through Thursday

Salary: $150,000–$250,000 base salary per year, plus equity and benefits

Company

hirify.global is an applied AI technology company that partners with major organizations to rearchitect critical operations and deploy AI-native systems for mission-critical workflows.

What you will do

  • Design and implement evaluation frameworks for AI systems deployed in customer environments.
  • Define domain-specific quality measurements that reflect user needs, operational constraints, and business objectives.
  • Build production Python test cases, golden datasets, regression suites, and AI-assisted test-generation workflows.
  • Develop offline and online evaluation pipelines that guide prompt design, agent logic, model selection, and release readiness.
  • Define and calibrate LLM-based graders against expert human assessments and investigate divergences between evaluation signals and real-world outcomes.
  • Collaborate with Forward Deployed AI Engineers, Architects, Product Engineers, AI Strategists, and domain experts on production system development and deployment.

Requirements

  • At least 2 years of software engineering experience.
  • Strong Python engineering skills and experience building maintainable production evaluation or experimentation pipelines.
  • Experience with evaluation-driven or experiment-driven development and understanding of metric overfitting risks.
  • Ability to translate subject-matter expertise and human judgment into test cases, scoring functions, and graders.
  • Systems-oriented understanding of how evaluation interacts with prompts, agents, data, and deployment.
  • Willingness to travel 10–50% depending on the project, role, and personal interest.

Culture & Benefits

  • 100% medical, dental, and vision insurance coverage for employees and dependents.
  • Flexible time off and retirement and financial planning benefits, including HSA, FSA, commuter accounts, 401(k), and financial coaching.
  • Wellness, mental well-being, fertility, and family-building benefits.
  • In-office lunches and snacks.
  • Access to modern AI models and tools, with ownership of high-impact projects for major enterprises.
  • Mission-driven, fast-moving culture focused on curiosity, pragmatism, and excellence.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →