Назад
Company hidden
1 день назад

Human Evaluation Researcher (AI)

160 000 - 190 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Human Evaluation Researcher (AI): Designing and running qualitative and quantitative studies to turn human judgments of real-time AI avatars into reliable evaluation signals with an accent on ambiguous judgments, emotional resonance, naturalness, and inter-rater agreement. Focus on building participant panels and evaluation pipelines, calibrating automated and LLM-based metrics against human ratings, and informing model training and release decisions.

Location: In-person in Seattle, five days a week

Base salary: $160,000–$190,000 per year, plus meaningful equity.

Company

hirify.global is a research company building photorealistic, real-time AI avatars with emotional intelligence and full-duplex audiovisual interaction.

What you will do

  • Design and run qualitative and quantitative studies of AI avatars, including side-by-side comparisons, controlled rating experiments, interviews, think-alouds, diary studies, and longitudinal panels.
  • Turn ambiguous judgments about naturalness, emotion, trust, and presence into aligned rubrics, anchored scales, and annotation guidelines.
  • Measure and improve inter-rater agreement while preserving meaningful human evaluation signal.
  • Use ethnographic methods, contextual inquiry, observation, and field work to understand real-world avatar experiences.
  • Build participant panels, rater training and calibration, evaluation tooling, and a cadence connected to model releases.
  • Calibrate automated and model-based metrics, including LLM-as-judge, against human judgment.

Requirements

  • 5+ years of human-subjects research experience in industry or academia, such as UX research, HCI, experimental psychology, or behavioral science.
  • Demonstrated experience designing studies that achieve alignment on ambiguous judgments such as tone, emotion, quality, or trust.
  • Strong qualitative research skills, including interviews, ethnography, and contextual inquiry.
  • Strong quantitative skills in survey and psychometric design, experimental design, and statistics for rating and pairwise-comparison data.
  • Fluency with agreement and reliability measures, including Cohen’s kappa and Krippendorff’s alpha.
  • Ability to run rigorous studies quickly and communicate findings clearly to ML researchers.

Nice to have

  • Experience evaluating generative AI, avatars, digital humans, speech or video generation, conversational agents, or emotional expression and recognition.
  • Background in perceptual science or psychophysics; an MS or PhD in a related field is welcome.
  • Experience with human evaluation at scale, crowdsourcing platforms, annotation tooling, or golden datasets.
  • Statistics and scripting experience with Python or R.

Culture & Benefits

  • Hands-on individual contributor role with end-to-end ownership of human evaluation.
  • Direct collaboration with the founders and modeling team.
  • Health plans including an HDHP with approximately $2,000 in annual employer HSA contributions.
  • 15 days of PTO, 10 public holidays, and a full week of office closure at year-end.
  • Workday meals, drinks, snacks, and commuter benefits of up to $340 per month.
  • 401(k) match and visa sponsorship, including O-1, H-1B, and green card support.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →