Назад
Company hidden
9 часов назад

Staff Slurm Cluster & HPC Engineer

Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Singapore/US/Norway +2 еще
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Staff Slurm Cluster & HPC Engineer (Slurm/Kubernetes/GPU Infrastructure): Designing and operating production Slurm clusters across bare-metal and virtualized GPU infrastructure with an accent on topology-aware scheduling, multi-tenant policy, and elastic capacity. Focus on integrating Slinky with Kubernetes, validating high-performance GPU fabrics, automating reproducible cluster delivery, and engineering reliable health checks, accounting, and customer-facing operations.

Location: Remote within the San Jose, CA or Austin, TX locations

Company

hirify.global provides Bitcoin mining infrastructure, AI computational infrastructure, data center operations, and cloud capabilities for high-demand artificial intelligence workloads.

What you will do

  • Design, deploy, and operate production Slurm clusters on bare-metal and virtual machine GPU infrastructure, including high availability, authentication, accounting, and live upgrades.
  • Build topology-aware scheduling for InfiniBand, RoCE, and NVLink GPU fabrics, validating placement quality through NCCL bandwidth and multi-node training tests.
  • Own multi-tenant scheduling policies covering accounts, associations, partitions, QOS, fairshare, preemption, reservations, and TRES limits with fail-closed authorization.
  • Lead Slinky slurm-operator adoption on Kubernetes and evaluate slurm-bridge for co-scheduling Kubernetes workloads through Slurm.
  • Automate elastic GPU capacity, containerized job runtimes, cluster health checks, rack acceptance testing, provisioning, observability, accounting, and billing integration.
  • Write runbooks and customer documentation, support enterprise onboarding and escalations, and mentor platform engineers.

Requirements

  • 8+ years of HPC, systems, or cloud infrastructure engineering experience, including 4+ years operating production Slurm clusters at 100+ GPU-node scale.
  • Deep hands-on experience with Slurm configuration, scheduling policies, accounting, authentication, REST APIs, and live version upgrades.
  • Strong knowledge of NVIDIA GPU infrastructure, DCGM, MIG, InfiniBand/RoCEv2, GPUDirect RDMA, NCCL tuning, and GPU failure diagnosis.
  • Production Kubernetes experience and hands-on exposure to Slinky slurm-operator, slurm-bridge, CoreWeave SUNK, or Nebius Soperator.
  • Experience with bare-metal provisioning, firmware lifecycle management, virtualized compute, Terraform, Ansible, and shared storage such as Lustre, GPFS, WEKA, VAST, or NFS.
  • Proficiency in Python and Bash, plus clear written and verbal English communication for enterprise customer engagement.

Nice to have

  • Go experience for platform control-plane and Slurm/Slinky REST integrations.

Culture & Benefits

  • Full-time role with direct ownership of Slurm architecture and HPC scheduling standards.
  • Hands-on work across bare-metal GPU infrastructure, virtualized clusters, Kubernetes, and AI workloads.
  • Direct collaboration with enterprise customers, product, sales, executive stakeholders, and platform engineers.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →