12 часов назад
Infrastructure Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Infrastructure Engineer (AI): Building and operating the research compute platform for a frontier AI lab, including GPU clusters, scheduling, networking, observability, and on-call systems with an accent on Kubernetes, Linux, cloud infrastructure, and large-scale AI workloads. Focus on optimizing GPU and CUDA infrastructure, designing agent-driven automation, and improving reliability and iteration speed for recursively self-improving AI research.
Location: London, United Kingdom; on-site
Company
is a well-funded, fast-growing frontier AI lab developing recursively self-improving AI to discover new scientific knowledge.
What you will do
- Run and evolve GPU clusters, including scheduling, utilization, debugging, and performance optimization.
- Scale the Kubernetes, Linux, networking, containers, and cloud infrastructure stack end to end.
- Establish operational excellence through incident response, postmortems, and healthy on-call practices.
- Build agent-driven automation for cluster lifecycle management, provisioning, and remediation.
- Partner directly with the research team to optimize the platform used for AI inference and training.
Requirements
- Experience operating infrastructure at scale, preferably for large language model workloads.
- Deep knowledge of Kubernetes internals, cluster provisioning, and orchestration systems.
- Practical expertise across Kubernetes, Linux, networking, containers, and cloud environments.
- Strong systems thinking with a focus on reliability, performance, and operational clarity.
- AI-focused mindset and interest in building agent-centered infrastructure.
Nice to have
- Cloud and cluster networking expertise, including VPC, BGP, CNI, eBPF, or service mesh.
- Experience with GPUs and CUDA.
- Infrastructure-as-code and workflow orchestration experience, such as Terraform.
- Experience leading multi-quarter infrastructure initiatives end to end.
Culture & Benefits
- Work in person every day at a high-intensity London headquarters.
- Shape the core technical foundation of a frontier AI lab from the beginning.
- Solve technically difficult infrastructure problems focused on accelerating AI research.
- Join a small, high-trust team with limited bureaucracy and a technical culture.
- Collaborate with experts in foundation model training, AI for science, and organizational design.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
2 часа назад
Senior Infrastructure Engineer (Kubernetes)
2 часа назад
Machine Learning & Cloud Infra Engineer (AI)
10 часов назад
Infrastructure Engineer (AI)
12 часов назад
Platform Engineer (AI)
3 часа назад
Principal Cloud Engineer (AI)
85 000 - 125 000GBP
6 дней назад
Infrastructure Operations Engineer (AI)
160 000 - 200 000$