Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Staff HPC Systems Software Engineer (AI/Slurm): Defining the architecture and technical direction of Slurm-based HPC services within a GPU cloud platform with an accent on cluster lifecycle, automation, reliability, and cloud-native integration. Focus on designing reusable platform patterns, integrating Slurm with Kubernetes-adjacent systems and infrastructure APIs, and solving complex GPU scheduling, HPC networking, and production supportability challenges.
Location: UK
Company
Nscale provides GPU cloud infrastructure for AI start-ups and enterprise customers, focusing on high-performance, cost-effective, and sustainable AI computing.
What you will do
- Define the technical direction and architecture for a core HPC systems domain, including Slurm platform architecture, scheduler integrations, cluster lifecycle, workload environments, or service automation.
- Package, automate, and expose Slurm implementations as services while establishing maintainable lifecycle and operating models.
- Lead cross-team design across Slurm, Kubernetes-adjacent systems, infrastructure APIs, identity systems, and platform tooling.
- Create reusable automation, deployment patterns, modules, and reference implementations to reduce duplicated effort and technical divergence.
- Lead technically critical initiatives across 2–4 teams and contribute hands-on engineering to de-risk complex work.
- Improve the reliability, observability, supportability, and maintainability of GPU-backed HPC services.
Requirements
- Extensive experience designing and building production software and automation for HPC systems, especially Slurm-based environments.
- Strong experience writing maintainable, testable, and resilient software in Go, Python, or similar languages.
- Deep understanding of Slurm internals, scheduler behavior, cluster lifecycle concerns, and operational trade-offs.
- Practical experience with GPU-backed infrastructure and HPC networking, including InfiniBand, RoCE, RDMA, and performance-sensitive workloads.
- Experience integrating HPC systems with cloud-native platforms, APIs, or service delivery models.
- Strong technical judgment, communication skills, and the ability to align multiple teams without formal authority.
Nice to have
- Experience with other schedulers or batch systems such as Kueue.
Culture & Benefits
- Opportunity to shape operating standards for a next-generation AI cloud platform.
- Work on complex infrastructure challenges with ownership and direct technical impact.
- Contribute to scaling high-performance and sustainable data center operations.
- Inclusive workplace with support for accessibility and individual accommodation needs.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
5 дней назад
Senior Software Engineer - AI/ML - AI Platform
9 дней назад
Member of Technical Staff - Backend (AI)
9 дней назад
Senior Infrastructure Software Engineer (AI)
180 000 - 220 000$
9 дней назад
Member of Technical Staff - AI
7 дней назад
Senior Software Engineer - Research Technology (C++)
Anthropic
7 дней назад
Staff Software Engineer, Observability & Profiling (AI)
325 000 - 390 000GBP