2 часа назад
Site Reliability Engineer (AI Infrastructure)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer (AI Infrastructure): Building and operating reliable platforms for large-scale GPU compute, Kubernetes, storage, and deployment systems with an accent on observability, automation, and distributed AI infrastructure. Focus on improving GPU node recovery, strengthening distributed training resilience, developing GitOps delivery, and optimizing capacity and resource efficiency.
Location: London, United Kingdom. Hybrid working model with in-person collaboration in the office and remote work.
Company
is building an AI platform for autonomous driving that enables vehicles to learn from real-world experience and adapt across different environments and vehicle platforms.
What you will do
- Build and operate platforms providing large-scale GPU compute, Kubernetes, storage, and deployment capabilities for AI development.
- Identify reliability risks and turn recurring operational issues into durable engineering solutions.
- Improve observability, automation, deployment safety, capacity planning, incident learning, SLOs, and production-readiness practices.
- Develop GPU node health detection and automated recovery, resilient distributed AI training, and continuous GitOps delivery with Argo CD or Flux.
- Enhance Kubernetes scheduling, autoscaling, policy enforcement, resource efficiency, and forecasting for GPU and storage capacity.
- Create reusable self-service capabilities and troubleshoot failures across compute, networking, storage, and distributed workloads.
Requirements
- Hands-on experience owning the reliability of large-scale cloud, Kubernetes, or distributed infrastructure.
- Recent experience coding, automating, configuring, or troubleshooting production systems.
- Proficiency in Python or Go, infrastructure as code, CI/CD or GitOps, and modern observability practices.
- Understanding of distributed-systems failure modes across compute, networking, and storage.
- Evidence of measurable improvements to reliability, performance, capacity, or cost, along with effective cross-team collaboration.
- Ability to work in a hybrid setup in London, United Kingdom.
Nice to have
- Direct experience with GPU, ML-training, or HPC infrastructure.
- Experience solving comparable large-scale platform challenges.
Culture & Benefits
- Hybrid working with core hours and opportunities to work hands-on in vehicle workshops and labs.
- Relocation support and visa sponsorship where applicable.
- Market-benchmarked salaries, meaningful equity, and location-dependent benefits.
- Learning and development budgets for training, conferences, and professional growth.
- Health insurance, dental coverage, enhanced parental leave, retirement or pension benefits where applicable, therapy access, wellbeing partnerships, and team socials.
Hiring process
- Initial recruiter call or screen.
- Competency interviews followed by deep-dive technical interviews.
- Final interview focused on mission and values alignment.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
Nebius
5 часов назад
Senior Site Reliability Engineer (AI/Cloud)
Nebius
5 часов назад
Senior Site Reliability Engineer (AI/ML Inference)
6 дней назад
Senior Site Reliability Engineer (AI)
4 дня назад
Senior Site Reliability Engineer (Kubernetes)
3 дня назад
Sr. Manager, Site Reliability
59 550 - 110 594GBP
6 дней назад
Senior Site Reliability Engineer (AI)
75 000 - 85 000$