Назад
обновлено 1 день назад

Staff HPC Systems Software Engineer (AI)

225 000 - 275 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Staff HPC Systems Software Engineer (Go/Python): Defining the technical direction and evolution of a Slurm-based HPC platform for a GPU cloud with an accent on cloud-native integration and automation. Focus on designing scalable cluster lifecycle models, optimizing GPU scheduling, and ensuring high-performance networking using InfiniBand and RDMA.

Location: New York, United States

Salary: $225,000–$275,000 USD annually base salary, plus potential bonus, equity, and/or commission.

Company

Nscale provides a GPU cloud infrastructure platform for AI startups and enterprise customers.

What you will do

  • Own the technical direction and architecture of a defined HPC systems domain, including Slurm platforms, scheduler integrations, cluster lifecycle, workload environments, or service automation.
  • Define how Slurm implementations are packaged, automated, and delivered as services across a cloud-native platform.
  • Establish shared patterns for automation, service lifecycle management, observability, reliability, and operational supportability.
  • Design integrations between Slurm, Kubernetes-adjacent systems, infrastructure APIs, identity systems, and platform tooling.
  • Create reusable modules, deployment patterns, automation, and reference implementations while reducing technical divergence and duplicated effort.
  • Lead critical initiatives across 2–4 teams, provide hands-on technical support, and influence engineering direction without formal authority.

Requirements

  • Extensive experience designing and building production software and automation for HPC systems, especially Slurm-based environments.
  • Strong software engineering experience with Go, Python, or similar languages, including maintainable, testable, and resilient software.
  • Deep understanding of Slurm internals, scheduler behavior, cluster lifecycle concerns, and operational trade-offs.
  • Practical experience with GPU-backed infrastructure and HPC networking, including InfiniBand, RoCE, RDMA, and performance-sensitive workloads.
  • Experience integrating HPC systems with cloud-native platforms, APIs, or service delivery models.
  • Ability to define technical direction, align multiple teams, and balance short-term delivery with long-term platform health.

Nice to have

  • Experience with other schedulers or batch systems such as Kueue.

Culture & Benefits

  • Collaborative, supportive, and innovation-focused working environment.
  • Flexible workplace with autonomy over daily scheduling.
  • Performance reviews every 12 months and a progression plan aligned with individual ambitions.
  • Potential medical, dental, vision, flexible paid time off, parental leave, and retirement plan benefits.
  • Opportunities to lead critical cross-functional initiatives and influence the planning and deployment of global AI capacity.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →