Назад
Company hidden
5 дней назад

Evaluation Research Manager (AI)

280 000 - 425 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
lead
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Evaluation Research Manager (AI): Leading rigorous evaluations of AI populations, predictions, and end-to-end simulations with an accent on experimental design, statistical measurement, and real-world validation. Focus on building reusable evaluation infrastructure, identifying failures hidden by aggregate metrics, and translating evidence into reliable product and research decisions.

Location: New York City; in-person five days a week in the office. Candidates must be located in the New York metropolitan area or open to relocation.

Salary: $280,000–$425,000 annually, plus equity.

Company

hirify.global builds simulations of human behavior using populations of AI agents to help companies and institutions test consequential decisions before acting.

What you will do

  • Lead and develop a team of Evaluation Researchers and research engineers.
  • Turn questions about realism, accuracy, calibration, usefulness, and decision quality into measurable constructs, experiments, and decision criteria.
  • Set the evaluation agenda across synthetic populations, predictive systems, agent behavior, group dynamics, and end-to-end simulations.
  • Build reusable evaluation infrastructure, including datasets, experiment harnesses, reporting systems, leaderboards, and protected holdouts.
  • Design studies using backtesting, prospective testing, temporal holdouts, subgroup analysis, uncertainty quantification, and reproducibility standards.
  • Communicate positive, negative, null, and inconclusive findings to research, engineering, product, customer, and company leadership audiences.

Requirements

  • Experience leading evaluation, measurement, or empirical research in machine learning, behavioral science, computational social science, statistics, economics, psychometrics, or a similarly rigorous field.
  • Experience building evaluations that influenced research direction, model capability, product decisions, or scientific conclusions.
  • Strong knowledge of experimental design, observational data, sampling, statistical power, uncertainty, causal threats, leakage, and condition shift.
  • Ability to write code, analyze large datasets, design studies, inspect individual failures, and review researchers' and engineers' technical work.
  • Experience managing or technically leading strong researchers and developing independent research judgment.
  • Willingness to work in person in New York City.

Nice to have

  • Experience with LLM agents, multi-agent systems, synthetic populations, recommender systems, probabilistic models, simulations, or decision-support tools.
  • Experience with longitudinal records, transaction data, product analytics, field experiments, prospective studies, backtesting, or operational outcomes.
  • Experience building evaluation platforms, regression suites, experiment-tracking systems, shared research datasets, model scorecards, or scientific reporting tools.
  • Experience communicating scientific results in customer-facing, public, policy, legal, or regulatory settings.

Culture & Benefits

  • Small, in-person team focused on urgency, high ownership, intellectual honesty, and truth-seeking.
  • Scientific independence is combined with close collaboration with research, engineering, product, deployment, and company leadership.
  • Comprehensive medical, vision, and dental coverage.
  • Equity participation, visa sponsorship, and relocation support.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →