Назад
Company hidden
2 часа назад

Members of Technical Staff (AI)

Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
UK
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Members of Technical Staff (AI) (Deep Learning/Post-Training): Developing post-training pipelines for foundation models that autonomously perform open-ended scientific research with an accent on SFT and RL algorithms, autocurricula, reward signals, and experimental analysis. Focus on building self-improving AI research agents, closing recursive research loops, and scaling multimodal model training on state-of-the-art hardware.

Location: London, United Kingdom; on-site, working in person every day

Company

hirify.global is a well-funded, fast-growing frontier AI research lab developing recursively self-improving AI for scientific discovery and AI-enabled science.

What you will do

  • Design, implement, and tune supervised fine-tuning and reinforcement learning algorithms for foundation models with open-ended scientific research capabilities.
  • Build autocurricula, judges, harnesses, and evaluation pipelines that convert open-ended research tasks into reliable reward signals.
  • Run large-scale experiments on state-of-the-art hardware and analyse results to define subsequent research hypotheses.
  • Develop recursive self-improvement loops in which AI agents contribute to their own post-training research.
  • Collaborate with Infrastructure and AI for Science teams to optimise hardware and deliver performance in scientific domains.

Requirements

  • 3+ years of deep learning research experience and 5+ years of software engineering experience.
  • Experience post-training large language, vision, video, or multimodal models.
  • Deep familiarity with Python and at least one deep learning framework, such as PyTorch or JAX.
  • Demonstrated deep learning research achievements through papers, model releases, open-source contributions, or comparable work.
  • Experience with modern coding agents and strong opinions on effective workflows.
  • On-site availability in London is required.

Nice to have

  • PhD in mathematics, computer science, or a hard science discipline.
  • Hands-on experience training large language models with reinforcement learning at scale, including GRPO, PPO, DPO, or distillation.
  • Familiarity with distributed and long-context training infrastructure.
  • Background in autocurricula, open-endedness, meta-learning, or recursive self-improvement.
  • Experience post-training frontier models at an industry lab.

Culture & Benefits

  • Opportunity to shape the core research of a frontier AI lab from its beginning.
  • Work on recursive self-improvement and AI scientists that improve the research pipeline that trains them.
  • Directly dogfood post-trained agents to accelerate further research.
  • Small, high-trust team with minimal bureaucracy and a strongly technical culture.
  • Emphasis on experimental organisational design, agent-centric workflows, and diversity of thought.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →