3 ΡΠ°ΡΠ° Π½Π°Π·Π°Π΄
Senior Research Scientist | Model Scaling (AI)
ΠΡΡΡ & Π‘ΠΎΠΏΡΠΎΠ²ΠΎΠ΄
ΠΠ»Ρ ΠΌΡΡΡΠ° Ρ ΡΡΠΎΠΉ Π²Π°ΠΊΠ°Π½ΡΠΈΠ΅ΠΉ Π½ΡΠΆΠ΅Π½ Plus
ΠΠΏΠΈΡΠ°Π½ΠΈΠ΅ Π²Π°ΠΊΠ°Π½ΡΠΈΠΈ
Π’Π΅ΠΊΡΡ:
TL;DR
Senior Research Scientist | Model Scaling (AI): Building and scaling next-generation language AI translation models with an accent on foundation-model selection, architecture decisions, and parameter-efficient adaptation. Focus on designing large-scale experiments, evaluating dense and Mixture-of-Experts models, and carrying research results through to production.
Location: Cologne, Germany; hybrid work with office attendance twice a week
Company
DeepL is a global AI product and research company building secure language AI solutions for translation, writing improvement, and real-time voice translation.
What you will do
- Select and evaluate open foundation and open-weight models for next-generation translation systems.
- Lead model selection and architecture decisions for scaling to hundreds of billions of parameters, including Mixture-of-Experts and other efficient designs.
- Design multi-capability adaptation strategies using LoRA, PEFT, and related methods.
- Own the modelling lifecycle from prototyping and ablation studies through scaling experiments, evaluation, and production delivery.
- Collaborate with post-training, RL/RLHF, and instruction-following specialists to integrate alignment and capabilities into base models.
- Track open-model and scaling research and translate findings into modelling recommendations.
Requirements
- Hands-on experience adapting and scaling multi-billion-parameter large language models through fine-tuning, instruction-tuning, or post-training.
- Strong understanding of architecture trade-offs at scale, including dense versus Mixture-of-Experts models.
- Working knowledge of parameter-efficient and multi-capability adaptation, including LoRA and PEFT.
- Experience training models, running experiments, debugging pipelines, and taking research results into production.
- Strong coding and experimentation skills with Python and PyTorch, JAX, or TensorFlow.
- Clear communication and effective collaboration across research, product, and engineering teams.
Nice to have
- Experience with uncertainty quantification, calibration, and confidence estimation for large models.
- Experience in machine translation, multilingual NLP, or document- and layout-aware modelling.
- Familiarity with MoE-specific training and adaptation, expert routing, Mixture-of-LoRA-Experts, and large-scale data-mixture design.
Culture & Benefits
- Internationally distributed team representing more than 90 nationalities.
- Open communication, regular feedback, and a collaborative growth-oriented environment.
- Hybrid schedule, flexible working hours, and coordination with team locations and time zones.
- Virtual Shares for every employee.
- Regular in-person team events and monthly full-day Hack Fridays.
- 30 days of annual leave excluding public holidays, mental health resources, and location-tailored benefits.
ΠΡΠ΄ΡΡΠ΅ ΠΎΡΡΠΎΡΠΎΠΆΠ½Ρ: Π΅ΡΠ»ΠΈ ΡΠ°Π±ΠΎΡΠΎΠ΄Π°ΡΠ΅Π»Ρ ΠΏΡΠΎΡΠΈΡ Π²ΠΎΠΉΡΠΈ Π² ΠΈΡ ΡΠΈΡΡΠ΅ΠΌΡ, ΠΈΡΠΏΠΎΠ»ΡΠ·ΡΡ iCloud/Google, ΠΏΡΠΈΡΠ»Π°ΡΡ ΠΊΠΎΠ΄/ΠΏΠ°ΡΠΎΠ»Ρ, Π·Π°ΠΏΡΡΡΠΈΡΡ ΠΊΠΎΠ΄/ΠΠ, Π½Π΅ Π΄Π΅Π»Π°ΠΉΡΠ΅ ΡΡΠΎΠ³ΠΎ - ΡΡΠΎ ΠΌΠΎΡΠ΅Π½Π½ΠΈΠΊΠΈ. ΠΠ±ΡΠ·Π°ΡΠ΅Π»ΡΠ½ΠΎ ΠΆΠΌΠΈΡΠ΅ "ΠΠΎΠΆΠ°Π»ΠΎΠ²Π°ΡΡΡΡ" ΠΈΠ»ΠΈ ΠΏΠΈΡΠΈΡΠ΅ Π² ΠΏΠΎΠ΄Π΄Π΅ΡΠΆΠΊΡ. ΠΠΎΠ΄ΡΠΎΠ±Π½Π΅Π΅ Π² Π³Π°ΠΉΠ΄Π΅ β