12 часов назад
Platform Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Platform Engineer (AI): Building and operating the infrastructure for large-scale reinforcement learning research, with an accent on GPU compute, Kubernetes, developer environments, and platform reliability. Focus on scheduling thousands of GPUs, diagnosing hardware and cluster-scale failures, and creating observability and tooling that accelerates experimentation.
Location: London, United Kingdom; on-site
Company
is developing a self-learning superlearner driven by large-scale reinforcement learning algorithms to discover knowledge and skills without relying on human data.
What you will do
- Design and own the platform infrastructure supporting large-scale reinforcement learning research.
- Manage Kubernetes clusters, containers, and internal tool deployments.
- Build GPU scheduling systems using tools such as KAI Scheduler and Kueue to keep thousands of GPUs productive.
- Improve infrastructure resilience by diagnosing hardware failures and preventing recurring cluster-scale issues.
- Develop observability through log management and monitoring with tools such as Quickwit, Grafana, or Datadog.
- Create fast, reliable developer environments and infrastructure tooling with Tailscale, Workbrew, dev containers, Python, and Rust.
Requirements
- Experience managing Kubernetes and containerized infrastructure.
- Hands-on experience with GPU scheduling at scale, including tools such as KAI Scheduler or Kueue.
- Experience operating resilient large-scale infrastructure and investigating hardware or cluster failures.
- Experience with Google Cloud or other major cloud providers.
- Ability to write clean infrastructure and platform tooling in Python and Rust.
Culture & Benefits
- Work as a foundational platform hire with meaningful influence over architecture and the shape of the role.
- Take ownership of infrastructure that directly supports ambitious AI research.
- Work in a fast-moving environment with trust, autonomy, and responsibility for engineering decisions.
- Collaborate on research infrastructure intended to advance reinforcement learning and artificial intelligence.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
16 часов назад
Platform Engineer (AI Infrastructure)
225 000 - 290 000$
12 часов назад
Infrastructure Engineer (AI)
3 часа назад
Principal Cloud Engineer (AI)
85 000 - 125 000GBP
12 часов назад
Platform Engineer (AI)
2 часа назад
Machine Learning & Cloud Infra Engineer (AI)
10 часов назад