Research Scientist (LLM Evaluations & Benchmarking)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
TL;DR
Research Scientist (LLM Evaluations & Benchmarking): Designing frontier-grade evaluations and benchmarks for reasoning, coding, agents, tool use, and multimodal models with an accent on construct validity, expert-verified ground truth, reliability, and contamination resistance. Focus on solving difficult measurement problems, leading expert calibration, validating results across models, and turning research into client pilots, public benchmarks, and papers.
Location: Fully remote from selected locations in Latin America and Europe, including Argentina, Brazil, Chile, Colombia, Ecuador, France, Mexico, Peru, the United Kingdom, and Uruguay.
Company
Labs develops AI evaluation and benchmarking solutions for frontier models and research laboratories.
What you will do
- Design original benchmarks for model reasoning, coding, agentic behavior, tool use, and multimodal capabilities.
- Build evaluation packages with subject-matter experts, expert-verified ground truth, multi-model results, calibration layers, weighted rubrics, and deterministic verifiers.
- Recruit, calibrate, and review expert pools across coding, agentic tool use, STEM, and reasoning.
- Act as a technical point of contact for AI labs and translate measurement goals into evaluation designs.
- Own pilots end to end and contribute to public benchmarks and research papers.
Requirements
- Research background in ML evaluation or benchmarking, demonstrated through published or open benchmarks, evaluation research, or equivalent hands-on work.
- Deep expertise in LLM and frontier-model benchmarking, especially code-model and agentic evaluation.
- Strong understanding of construct validity, psychometrics, rubrics, pass rates, headroom, contamination, and task discrimination.
- Interest in safety evaluation, capability elicitation, and robustness research.
- Ability to maintain rigorous standards across a research team or expert pool and complete the full research cycle from framing to publication.
- Fluent English required; Spanish is a plus.
Nice to have
- Spanish language proficiency.
- Experience publishing at NeurIPS Datasets & Benchmarks, ICLR, ACL, or similar venues.
Culture & Benefits
- Fully remote contract role within the listed countries.
- Work directly with frontier AI labs and expert research communities.
- Opportunity to develop public benchmarks and publish research.
- Close collaboration with the CEO and internal AI research team.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →