Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Staff Cloud Native Software Engineer (AI/Kubernetes): Building and operating Kubernetes-native controllers, operators, CRDs, and networking integrations that connect AI applications with GPU-backed infrastructure with an accent on reliability, scalability, and observability. Focus on extending Kubernetes control-plane capabilities, designing self-healing reconciliation systems, debugging CNI and service-mesh integrations, and setting technical direction across teams.
Location: UK
Company
Nscale provides GPU cloud infrastructure for AI start-ups and enterprise customers, focused on high-performance, cost-effective, and sustainable AI computing.
What you will do
- Design, build, and operate Kubernetes-native controllers, operators, custom resources, and admission webhooks for AI applications and networking components.
- Extend Kubernetes control-plane capabilities for AI workloads through network policy controllers, CNI and service-mesh integrations, and resource and scheduling extensions.
- Build reconciliation loops, informers, client-go tooling, and operational automation that keep infrastructure state consistent and simplify platform operations.
- Drive infrastructure architecture decisions and build observability foundations using metrics, structured events, tracing, and Kubernetes status reporting.
- Debug complex issues across Kubernetes, networking datapaths, service meshes, and GPU-backed workload runtimes while designing systems that degrade gracefully and self-heal.
- Set technical direction, lead design discussions and code reviews, define reusable platform patterns, and mentor engineers.
Requirements
- At least 8 years of production-level software development experience.
- Deep hands-on experience building and operating Kubernetes-native software, including controllers, operators, CRDs, or admission webhooks.
- Strong understanding of Kubernetes internals, including the API server, informer and lister patterns, reconciliation loops, and the object model.
- Strong networking fundamentals covering CNI, service mesh, kube-proxy or eBPF datapaths, DNS, and load balancing.
- Proficiency in Go or a similar language, with experience delivering tested, production-quality software at scale.
- Experience with observability, independent end-to-end ownership, staff-level technical leadership, and mentoring.
Nice to have
- Experience with or strong interest in GPU-backed infrastructure and AI workload patterns.
Culture & Benefits
- Opportunity to shape operating standards for a next-generation AI cloud platform.
- Work on complex infrastructure challenges involving high-performance and sustainable data centre operations.
- Culture centered on innovation, ownership, accountability, openness, and transparency.
- Inclusive workplace with accommodations available for individual needs.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
5 дней назад
Infrastructure Software Engineer, Apps Platform
5 дней назад
Staff Software Engineer (Customer Platform)
Anthropic
5 дней назад
Staff Software Engineer, Observability & Profiling (AI)
325 000 - 390 000GBP
Nebius
4 дня назад
Senior/Tech Lead Backend Engineer (Managed Kubernetes)
5 дней назад
Senior Software Engineer (AI CICD)
157 000 - 184 000$
9 часов назад