3 дня назад
Infrastructure Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Infrastructure Engineer (AI) (distributed training and inference systems): Building the training, inference, orchestration, and data pipeline infrastructure that frontier AI research depends on with an accent on speed, reliability, large-scale distributed RL, and scientific deployment. Focus on hardening RL and post-training libraries, profiling performance, and closing recursive loops so AI agents can drive their own training and infrastructure.
Location: London, United Kingdom; on-site and in-person every day
Company
is a well-funded, fast-growing frontier AI lab developing recursively self-improving AI to discover new scientific knowledge.
What you will do
- Contribute end-to-end to training and inference infrastructure for frontier research.
- Build distributed job orchestration, data pipelines, evaluation systems, and infrastructure for large-scale experiments.
- Implement and harden reinforcement learning and post-training libraries for AI models.
- Develop systems that enable AI agents to drive their own training and infrastructure.
- Collaborate closely with the core research team on the systems used for model development and scientific research.
Requirements
- Experience building high-performance, large-scale distributed systems, preferably for LLM workloads.
- Proficiency in Python and PyTorch or JAX.
- Hands-on experience with large-scale LLM training or inference technologies such as SGLang, vLLM, verl, Megatron, or OpenRLHF.
- Experience with a systems programming language such as Rust or C++.
- Experience writing and profiling CUDA kernels.
- A track record of building reliable research tools and a strong judgment about when to build, buy, or delete.
Nice to have
- Experience with CUDA kernel development and profiling.
- Experience building reliable tools for research environments.
Culture & Benefits
- Shape the core technical foundation of a frontier AI lab from the beginning.
- Work on unusually difficult and creative infrastructure problems involving recursively self-improving agents.
- Join a small, high-trust team with no bureaucracy and a strongly technical culture.
- Collaborate with experts in foundation model training, AI for science, and organisational design.
- Unusual career paths and diverse backgrounds are welcomed.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →