Эта вакансия в архиве

Посмотреть похожие вакансии ↓
Company hidden
5 дней назад

Senior Slurm Cluster & HPC Engineer

Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Singapore

Описание вакансии

Текст:
/
TL;DR
Senior Slurm Cluster & HPC Engineer (Slurm/GPU/Kubernetes): Designing, operating, and automating production Slurm clusters and GPU cloud capacity across bare metal, virtualized infrastructure, and Kubernetes with an accent on topology-aware scheduling, multi-tenant isolation, and cluster reliability. Focus on integrating Slurm with Slinky, validating InfiniBand/RoCE and NCCL performance, automating elastic GPU capacity, and detecting failures with automatic drain and job requeue.

Location: Singapore, SG

Company

hirify.global is a technology company providing Bitcoin mining solutions and AI cloud infrastructure, including ASIC manufacturing, datacenter operations, and high-performance computing services.

What you will do

  • Design, deploy, and operate production Slurm clusters on bare metal and virtual machines, including high availability, authentication, accounting, and live version upgrades.
  • Build topology-aware GPU scheduling for InfiniBand, RoCE, and NVLink fabrics, validating placement with NCCL bandwidth and multi-node training tests.
  • Own multi-tenant scheduling policies covering accounts, associations, partitions, QOS, fair share, preemption, reservations, and tenant resource limits.
  • Lead Slurm-on-Kubernetes delivery with Slinky, evaluate slurm-bridge, and manage elastic GPU capacity between Slurm batch workloads and Kubernetes inference.
  • Operate containerized job runtimes, shared storage, MPI/PMIx environments, and automated health checks with GPU, network, and hardware fault detection.
  • Deliver reproducible infrastructure through Terraform, Ansible, PXE, Redfish, and IPMI; integrate observability, accounting, billing, customer support, and technical mentoring.

Requirements

  • 8+ years of experience in HPC, systems, or cloud infrastructure engineering, including 4+ years operating production Slurm clusters at 100+ GPU-node scale.
  • Deep hands-on knowledge of Slurm configuration, scheduling policy, accounting, authentication, REST services, and live cluster upgrades.
  • Strong experience with NVIDIA GPU infrastructure, DCGM, MIG, InfiniBand/RoCEv2, GPUDirect RDMA, NCCL, and GPU failure diagnosis.
  • Production Kubernetes experience and exposure to Slinky slurm-operator, slurm-bridge, or a comparable Slurm-on-Kubernetes stack.
  • Experience with bare-metal provisioning, firmware and BIOS lifecycle management, virtualized compute, Terraform, Ansible, and parallel or shared storage.
  • Proficiency in Python and Bash, strong multi-tenant security practices, and clear written and verbal English for enterprise customer engagement.

Nice to have

  • Go experience for platform control-plane and Slurm/Slinky REST integrations.

Culture & Benefits

  • Inclusive environment valuing authenticity and diverse perspectives.
  • Startup-oriented atmosphere with autonomy, accountability, and opportunities for rapid growth.
  • Direct contribution to digital asset and AI infrastructure projects.
  • Training, mentoring, welfare benefits, and professional development opportunities.