Назад
Company hidden
обновлено 5 дней назад

Research Engineer (AI Safety)

Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
c1
Страна
France/UK
Релокация
France
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Research Engineer (AI Safety): Building and maintaining an internal suite of benchmarks for content and agentic guardrails with an accent on model capabilities and agent behaviors. Focus on quantifying realistic LLM failure modes in the wild and creating defensible capability benchmarks.

Location: Hybrid work from the Paris or London offices. Relocation support is available for moves to Paris after the probationary period.

Company

hirify.global is an AI Safety company building a safety, reliability, and optimization layer for AI systems through natural-language policies, automated testing, enforcement, and continuous improvement.

What you will do

  • Own and maintain internal benchmarks for single- and multi-turn content guardrails and agentic safety.
  • Build benchmarks that distinguish specific model capabilities with measurable and defensible results.
  • Work with the product team to create evaluations for flagship model functionality.
  • Develop benchmarks for new research features and adapt evaluations to new verticals and changing product data.
  • Study and quantify realistic agentic and LLM failure modes in real-world settings.

Requirements

  • Experience building an LLM benchmark from scratch that distinguishes specific model capabilities.
  • Experience creating synthetic data for post-training textual or multimodal models.
  • Ability to reproduce published benchmark results and identify fragile or misleading methodology.
  • Production Python experience, including writing maintainable code and designing extensible abstractions.
  • Ability to build efficient LLM inference setups with parallel-call orchestration, retries, and rate-limit handling.
  • Strong practical familiarity with frontier models and coding agents.

Nice to have

  • Automated red-teaming experience.
  • Experience with agentic scaffolds and public benchmark reproduction.
  • Knowledge of reward-model, monitoring, and safety benchmarks.
  • Published papers in evaluations or safety evaluation.

Culture & Benefits

  • Small, focused team working on AI safety research and production systems.
  • Competitive compensation including equity.
  • Flexible time off and flexible hybrid work.
  • Premium private health insurance and mental health support, including therapy coverage.
  • Office meals, learning and development support, and required hardware and software.
  • Team off-sites twice a year.

Hiring process

  • Introductory call with the Talent Team.
  • Test assignment.
  • Technical interview with the Head of Fundamental Research, followed by a final interview with the CEO.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →