2 часа назад
Members of Technical Staff (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Members of Technical Staff (AI) (Deep Learning/Post-Training): Developing post-training pipelines for foundation models that autonomously perform open-ended scientific research with an accent on SFT and RL algorithms, autocurricula, reward signals, and experimental analysis. Focus on building self-improving AI research agents, closing recursive research loops, and scaling multimodal model training on state-of-the-art hardware.
Location: London, United Kingdom; on-site, working in person every day
Company
is a well-funded, fast-growing frontier AI research lab developing recursively self-improving AI for scientific discovery and AI-enabled science.
What you will do
- Design, implement, and tune supervised fine-tuning and reinforcement learning algorithms for foundation models with open-ended scientific research capabilities.
- Build autocurricula, judges, harnesses, and evaluation pipelines that convert open-ended research tasks into reliable reward signals.
- Run large-scale experiments on state-of-the-art hardware and analyse results to define subsequent research hypotheses.
- Develop recursive self-improvement loops in which AI agents contribute to their own post-training research.
- Collaborate with Infrastructure and AI for Science teams to optimise hardware and deliver performance in scientific domains.
Requirements
- 3+ years of deep learning research experience and 5+ years of software engineering experience.
- Experience post-training large language, vision, video, or multimodal models.
- Deep familiarity with Python and at least one deep learning framework, such as PyTorch or JAX.
- Demonstrated deep learning research achievements through papers, model releases, open-source contributions, or comparable work.
- Experience with modern coding agents and strong opinions on effective workflows.
- On-site availability in London is required.
Nice to have
- PhD in mathematics, computer science, or a hard science discipline.
- Hands-on experience training large language models with reinforcement learning at scale, including GRPO, PPO, DPO, or distillation.
- Familiarity with distributed and long-context training infrastructure.
- Background in autocurricula, open-endedness, meta-learning, or recursive self-improvement.
- Experience post-training frontier models at an industry lab.
Culture & Benefits
- Opportunity to shape the core research of a frontier AI lab from its beginning.
- Work on recursive self-improvement and AI scientists that improve the research pipeline that trains them.
- Directly dogfood post-trained agents to accelerate further research.
- Small, high-trust team with minimal bureaucracy and a strongly technical culture.
- Emphasis on experimental organisational design, agent-centric workflows, and diversity of thought.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
5 дней назад
Technical Lead, Machine Learning (AI)
4 часа назад
Senior Member of Technical Staff, Machine Learning (AI)
1 день назад
Research Scientist/Research Engineer, Reinforcement Learning
200 000 - 350 000$
3 часа назад
AI Research Residency (Agentic AI)
150 000$
DeepL
7 дней назад
Senior Staff Research Scientist (AI)
1 день назад