Назад
7 дней назад

Staff HPC Systems Software Engineer (AI)

Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
UK
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Staff HPC Systems Software Engineer (AI/Slurm): Defining the architecture and technical direction of Slurm-based HPC services within a GPU cloud platform with an accent on cluster lifecycle, automation, reliability, and cloud-native integration. Focus on designing reusable platform patterns, integrating Slurm with Kubernetes-adjacent systems and infrastructure APIs, and solving complex GPU scheduling, HPC networking, and production supportability challenges.

Location: UK

Company

Nscale provides GPU cloud infrastructure for AI start-ups and enterprise customers, focusing on high-performance, cost-effective, and sustainable AI computing.

What you will do

  • Define the technical direction and architecture for a core HPC systems domain, including Slurm platform architecture, scheduler integrations, cluster lifecycle, workload environments, or service automation.
  • Package, automate, and expose Slurm implementations as services while establishing maintainable lifecycle and operating models.
  • Lead cross-team design across Slurm, Kubernetes-adjacent systems, infrastructure APIs, identity systems, and platform tooling.
  • Create reusable automation, deployment patterns, modules, and reference implementations to reduce duplicated effort and technical divergence.
  • Lead technically critical initiatives across 2–4 teams and contribute hands-on engineering to de-risk complex work.
  • Improve the reliability, observability, supportability, and maintainability of GPU-backed HPC services.

Requirements

  • Extensive experience designing and building production software and automation for HPC systems, especially Slurm-based environments.
  • Strong experience writing maintainable, testable, and resilient software in Go, Python, or similar languages.
  • Deep understanding of Slurm internals, scheduler behavior, cluster lifecycle concerns, and operational trade-offs.
  • Practical experience with GPU-backed infrastructure and HPC networking, including InfiniBand, RoCE, RDMA, and performance-sensitive workloads.
  • Experience integrating HPC systems with cloud-native platforms, APIs, or service delivery models.
  • Strong technical judgment, communication skills, and the ability to align multiple teams without formal authority.

Nice to have

  • Experience with other schedulers or batch systems such as Kueue.

Culture & Benefits

  • Opportunity to shape operating standards for a next-generation AI cloud platform.
  • Work on complex infrastructure challenges with ownership and direct technical impact.
  • Contribute to scaling high-performance and sustainable data center operations.
  • Inclusive workplace with support for accessibility and individual accommodation needs.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →