9 часов назад
Staff Slurm Cluster & HPC Engineer
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Staff Slurm Cluster & HPC Engineer (Slurm/Kubernetes/GPU Infrastructure): Designing and operating production Slurm clusters across bare-metal and virtualized GPU infrastructure with an accent on topology-aware scheduling, multi-tenant policy, and elastic capacity. Focus on integrating Slinky with Kubernetes, validating high-performance GPU fabrics, automating reproducible cluster delivery, and engineering reliable health checks, accounting, and customer-facing operations.
Location: Remote within the San Jose, CA or Austin, TX locations
Company
provides Bitcoin mining infrastructure, AI computational infrastructure, data center operations, and cloud capabilities for high-demand artificial intelligence workloads.
What you will do
- Design, deploy, and operate production Slurm clusters on bare-metal and virtual machine GPU infrastructure, including high availability, authentication, accounting, and live upgrades.
- Build topology-aware scheduling for InfiniBand, RoCE, and NVLink GPU fabrics, validating placement quality through NCCL bandwidth and multi-node training tests.
- Own multi-tenant scheduling policies covering accounts, associations, partitions, QOS, fairshare, preemption, reservations, and TRES limits with fail-closed authorization.
- Lead Slinky slurm-operator adoption on Kubernetes and evaluate slurm-bridge for co-scheduling Kubernetes workloads through Slurm.
- Automate elastic GPU capacity, containerized job runtimes, cluster health checks, rack acceptance testing, provisioning, observability, accounting, and billing integration.
- Write runbooks and customer documentation, support enterprise onboarding and escalations, and mentor platform engineers.
Requirements
- 8+ years of HPC, systems, or cloud infrastructure engineering experience, including 4+ years operating production Slurm clusters at 100+ GPU-node scale.
- Deep hands-on experience with Slurm configuration, scheduling policies, accounting, authentication, REST APIs, and live version upgrades.
- Strong knowledge of NVIDIA GPU infrastructure, DCGM, MIG, InfiniBand/RoCEv2, GPUDirect RDMA, NCCL tuning, and GPU failure diagnosis.
- Production Kubernetes experience and hands-on exposure to Slinky slurm-operator, slurm-bridge, CoreWeave SUNK, or Nebius Soperator.
- Experience with bare-metal provisioning, firmware lifecycle management, virtualized compute, Terraform, Ansible, and shared storage such as Lustre, GPFS, WEKA, VAST, or NFS.
- Proficiency in Python and Bash, plus clear written and verbal English communication for enterprise customer engagement.
Nice to have
- Go experience for platform control-plane and Slurm/Slinky REST integrations.
Culture & Benefits
- Full-time role with direct ownership of Slurm architecture and HPC scheduling standards.
- Hands-on work across bare-metal GPU infrastructure, virtualized clusters, Kubernetes, and AI workloads.
- Direct collaboration with enterprise customers, product, sales, executive stakeholders, and platform engineers.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
15 часов назад
Senior HPC Cluster Engineer
145 920 - 209 241$
Lambda
5 дней назад
Site Reliability Engineer (AI Infrastructure)
240 000 - 356 000$
17 часов назад
HPC Engineer (AI)
2 дня назад
Data Center Provisioning Engineer
12 часов назад
AI Infrastructure Solutions Engineer
8 часов назад
Cloud Infrastructure Engineer
150 000 - 167 166$