обновлено 1 день назад
Staff HPC Systems Software Engineer (AI)
225 000 - 275 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Staff HPC Systems Software Engineer (Go/Python): Defining the technical direction and evolution of a Slurm-based HPC platform for a GPU cloud with an accent on cloud-native integration and automation. Focus on designing scalable cluster lifecycle models, optimizing GPU scheduling, and ensuring high-performance networking using InfiniBand and RDMA.
Location: New York, United States
Salary: $225,000–$275,000 USD annually base salary, plus potential bonus, equity, and/or commission.
Company
Nscale provides a GPU cloud infrastructure platform for AI startups and enterprise customers.
What you will do
- Own the technical direction and architecture of a defined HPC systems domain, including Slurm platforms, scheduler integrations, cluster lifecycle, workload environments, or service automation.
- Define how Slurm implementations are packaged, automated, and delivered as services across a cloud-native platform.
- Establish shared patterns for automation, service lifecycle management, observability, reliability, and operational supportability.
- Design integrations between Slurm, Kubernetes-adjacent systems, infrastructure APIs, identity systems, and platform tooling.
- Create reusable modules, deployment patterns, automation, and reference implementations while reducing technical divergence and duplicated effort.
- Lead critical initiatives across 2–4 teams, provide hands-on technical support, and influence engineering direction without formal authority.
Requirements
- Extensive experience designing and building production software and automation for HPC systems, especially Slurm-based environments.
- Strong software engineering experience with Go, Python, or similar languages, including maintainable, testable, and resilient software.
- Deep understanding of Slurm internals, scheduler behavior, cluster lifecycle concerns, and operational trade-offs.
- Practical experience with GPU-backed infrastructure and HPC networking, including InfiniBand, RoCE, RDMA, and performance-sensitive workloads.
- Experience integrating HPC systems with cloud-native platforms, APIs, or service delivery models.
- Ability to define technical direction, align multiple teams, and balance short-term delivery with long-term platform health.
Nice to have
- Experience with other schedulers or batch systems such as Kueue.
Culture & Benefits
- Collaborative, supportive, and innovation-focused working environment.
- Flexible workplace with autonomy over daily scheduling.
- Performance reviews every 12 months and a progression plan aligned with individual ambitions.
- Potential medical, dental, vision, flexible paid time off, parental leave, and retirement plan benefits.
- Opportunities to lead critical cross-functional initiatives and influence the planning and deployment of global AI capacity.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
6 дней назад
Cloud Systems Engineer (AI/HPC)
100 000 - 135 000$
Microsoft AI
6 дней назад
Member of Technical Staff - AI Networking
119 800 - 234 700$
6 дней назад
High Performance Compute Systems Site Lead (Onsite - LANL)
105 500 - 243 000$
Anthropic
5 дней назад
Staff Engineer, Datacenter Server Lifecycle
320 000 - 405 000$
8 дней назад
High Performance Computing (HPC) AI Engineer
101 494 - 140 000$
3 дня назад
Senior Data Center Technician (AI/HPC)
76 000 - 94 000$