Назад
Company hidden
1 день назад

Research Scientist (LLM Evaluations & Benchmarking)

Формат работы
remote (только Europe)
Тип работы
project
Английский
c1
Страна
France/UK/US +8 еще
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/

TL;DR

Research Scientist (LLM Evaluations & Benchmarking): Designing frontier-grade evaluations and benchmarks for reasoning, coding, agents, tool use, and multimodal models with an accent on construct validity, expert-verified ground truth, reliability, and contamination resistance. Focus on solving difficult measurement problems, leading expert calibration, validating results across models, and turning research into client pilots, public benchmarks, and papers.

Location: Fully remote from selected locations in Latin America and Europe, including Argentina, Brazil, Chile, Colombia, Ecuador, France, Mexico, Peru, the United Kingdom, and Uruguay.

Company

hirify.global Labs develops AI evaluation and benchmarking solutions for frontier models and research laboratories.

What you will do

  • Design original benchmarks for model reasoning, coding, agentic behavior, tool use, and multimodal capabilities.
  • Build evaluation packages with subject-matter experts, expert-verified ground truth, multi-model results, calibration layers, weighted rubrics, and deterministic verifiers.
  • Recruit, calibrate, and review expert pools across coding, agentic tool use, STEM, and reasoning.
  • Act as a technical point of contact for AI labs and translate measurement goals into evaluation designs.
  • Own pilots end to end and contribute to public benchmarks and research papers.

Requirements

  • Research background in ML evaluation or benchmarking, demonstrated through published or open benchmarks, evaluation research, or equivalent hands-on work.
  • Deep expertise in LLM and frontier-model benchmarking, especially code-model and agentic evaluation.
  • Strong understanding of construct validity, psychometrics, rubrics, pass rates, headroom, contamination, and task discrimination.
  • Interest in safety evaluation, capability elicitation, and robustness research.
  • Ability to maintain rigorous standards across a research team or expert pool and complete the full research cycle from framing to publication.
  • Fluent English required; Spanish is a plus.

Nice to have

  • Spanish language proficiency.
  • Experience publishing at NeurIPS Datasets & Benchmarks, ICLR, ACL, or similar venues.

Culture & Benefits

  • Fully remote contract role within the listed countries.
  • Work directly with frontier AI labs and expert research communities.
  • Opportunity to develop public benchmarks and publish research.
  • Close collaboration with the CEO and internal AI research team.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →