Назад
26 дней назад

Senior Research Scientist (LLM Post-Training)

Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
UK
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Research Scientist (LLM Post-Training): Lead fine-tuning, post-training, model-steerability, and reinforcement learning for DeepL’s next-generation LLM-based translation models with an accent on steerable instruction-following translation and large-scale experimentation. Focus on building reward/evaluator models, mitigating reward hacking and quality-estimation failures, and shipping improvements into production real-time systems.

Location: London

Company

DeepL is a global AI product and research company building secure language AI for translation, writing, and voice translation.

What you will do

  • Develop steerable translation models conditioned on user preferences, rules, and context.
  • Conduct hands-on post-training research for core translation models, including supervised fine-tuning, knowledge distillation, preference optimization, and translation-quality-tuned reinforcement learning.
  • Build reward models and evaluator models for translation, including rubric- and reference-based grading, and investigate/mitigate reward hacking and quality-estimation failure modes.
  • Advance multimodal translation by driving research toward models that ingest multimodal content and context.
  • Own the full model lifecycle from prototyping and ablations to training, evaluation, optimization, and production deployment with engineering.
  • Set up evaluation, reproducibility, monitoring, and continuous improvement practices; mentor researchers and engineers.

Requirements

  • Proven experience making large models steerable and instruction-following using effective behavior-instillation methods (e.g., instruction tuning, latent space methods, steering vectors, constrained encoding/decoding).
  • Deep hands-on expertise in LLM post-training (SFT, DPO), knowledge distillation, and/or reinforcement learning (RLHF/RLAIF, PPO/GSPO, reward modeling).
  • Strong data-centric skills for synthetic-data and preference-data pipelines, including LLM-as-judge generation, data curation/filtering, and reasoning about data mixtures and ablations.
  • Experience designing evaluation and reward signals using automatic metrics, LLM-as-judge evaluation, non-verifiable rewards, and human-in-the-loop evaluation.
  • Hands-on builder mindset: training models, running experiments, debugging pipelines, and integrating ML systems into production.
  • Strong coding and experimentation skills (Python, PyTorch/JAX/TensorFlow) and clear communication aligning research with product and engineering priorities.

Culture & Benefits

  • Hybrid work schedule with office attendance twice a week.
  • Flexible working hours aligned with team time zones.
  • Virtual Shares for an ownership mindset.
  • Regular in-person team events and monthly full-day hacking sessions (Hack Fridays).
  • 30 days of annual leave (excluding public holidays) and access to mental health resources.
  • Competitive benefits tailored to location.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →