Назад
Company hidden
3 часа назад

Member of Technical Staff, Model Evaluations (AI)

200 000 - 400 000$
Формат работы
onsite
Тип работы
fulltime
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Member of Technical Staff, Model Evaluations (AI): Building measurement systems, evaluation suites, metrics, rubrics, datasets, dashboards, and workflows to determine whether behavioral simulations are accurate and trustworthy, with an accent on statistical judgment, human data, and model quality. Focus on evaluating LLM and product outputs, making uncertainty and ground truth legible, and automating rigorous evaluation workflows for real-world decision-making.

Location: On-site in San Francisco or New York, United States

Salary: $200,000–$400,000 USD annually, plus equity and comprehensive benefits.

Company

hirify.global builds AI infrastructure for simulating human behavior and helping organizations make evidence-based decisions.

What you will do

  • Design evaluation systems, metrics, rubrics, datasets, dashboards, and workflows for behavioral simulation models.
  • Evaluate model versions, diagnose regressions, identify improvement priorities, and maintain stable evaluation suites.
  • Build evaluations for qualitative responses, retrieval, survey generation, AI-generated research reports, and other customer-facing outputs.
  • Compare simulations with human data, customer studies, collected ground truth, behavioral datasets, and real-world signals.
  • Develop internal tools, labeling workflows, automated graders, and evaluation pipelines using agentic coding tools.
  • Prototype evaluation methods for transaction behavior, product interactions, experiments, interventions, and multi-agent group settings.

Requirements

  • Strong judgment about whether evaluations are meaningful, robust, measurable, and relevant to product or model decisions.
  • Working knowledge of LLM training, post-training, model evaluation, and model improvement cycles.
  • Comfort with noisy data, uncertainty, sampling, distributions, calibration, confidence intervals, bias, variance, and measurement validity.
  • Ability to build tools, scripts, dashboards, analyses, labeling workflows, or automated evaluation pipelines.
  • Experience with data and automation tools such as Python, SQL, R, notebooks, LLM APIs, and agentic coding tools.
  • Ability to work on-site in San Francisco or New York.

Nice to have

  • Experience with model evaluation dashboards, regression suites, release gates, benchmarks, or model comparison workflows.
  • Experience with LLM-as-judge systems, human data, rubrics, automated graders, expert reviews, or grader calibration.
  • Background in survey methodology, behavioral science, experimentation, market research, polling, UXR, or computational social science.
  • Experience evaluating behavioral signals such as transactions, purchases, mobility, or product interactions.
  • Interest or experience in multi-agent behavior, group conversation, deliberation, or collective decision-making.

Culture & Benefits

  • Work with a research and engineering team building behavioral simulation infrastructure.
  • Competitive compensation including base salary and equity for eligible roles.
  • Comprehensive medical, dental, and vision coverage.
  • Flexible time-off policies.

Hiring process

  • Hiring conversations focus on fit, working style, expectations, and clear examples of past work.
  • Candidates may reapply for the same role after a 90-day waiting period.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →