Эта вакансия в архиве
Посмотреть похожие вакансии ↓5 дней назад
Senior Slurm Cluster & HPC Engineer
Описание вакансии
Текст:
TL;DR
Senior Slurm Cluster & HPC Engineer (Slurm/GPU/Kubernetes): Designing, operating, and automating production Slurm clusters and GPU cloud capacity across bare metal, virtualized infrastructure, and Kubernetes with an accent on topology-aware scheduling, multi-tenant isolation, and cluster reliability. Focus on integrating Slurm with Slinky, validating InfiniBand/RoCE and NCCL performance, automating elastic GPU capacity, and detecting failures with automatic drain and job requeue.
Location: Singapore, SG
Company
is a technology company providing Bitcoin mining solutions and AI cloud infrastructure, including ASIC manufacturing, datacenter operations, and high-performance computing services.
What you will do
- Design, deploy, and operate production Slurm clusters on bare metal and virtual machines, including high availability, authentication, accounting, and live version upgrades.
- Build topology-aware GPU scheduling for InfiniBand, RoCE, and NVLink fabrics, validating placement with NCCL bandwidth and multi-node training tests.
- Own multi-tenant scheduling policies covering accounts, associations, partitions, QOS, fair share, preemption, reservations, and tenant resource limits.
- Lead Slurm-on-Kubernetes delivery with Slinky, evaluate slurm-bridge, and manage elastic GPU capacity between Slurm batch workloads and Kubernetes inference.
- Operate containerized job runtimes, shared storage, MPI/PMIx environments, and automated health checks with GPU, network, and hardware fault detection.
- Deliver reproducible infrastructure through Terraform, Ansible, PXE, Redfish, and IPMI; integrate observability, accounting, billing, customer support, and technical mentoring.
Requirements
- 8+ years of experience in HPC, systems, or cloud infrastructure engineering, including 4+ years operating production Slurm clusters at 100+ GPU-node scale.
- Deep hands-on knowledge of Slurm configuration, scheduling policy, accounting, authentication, REST services, and live cluster upgrades.
- Strong experience with NVIDIA GPU infrastructure, DCGM, MIG, InfiniBand/RoCEv2, GPUDirect RDMA, NCCL, and GPU failure diagnosis.
- Production Kubernetes experience and exposure to Slinky slurm-operator, slurm-bridge, or a comparable Slurm-on-Kubernetes stack.
- Experience with bare-metal provisioning, firmware and BIOS lifecycle management, virtualized compute, Terraform, Ansible, and parallel or shared storage.
- Proficiency in Python and Bash, strong multi-tenant security practices, and clear written and verbal English for enterprise customer engagement.
Nice to have
- Go experience for platform control-plane and Slurm/Slinky REST integrations.
Culture & Benefits
- Inclusive environment valuing authenticity and diverse perspectives.
- Startup-oriented atmosphere with autonomy, accountability, and opportunities for rapid growth.
- Direct contribution to digital asset and AI infrastructure projects.
- Training, mentoring, welfare benefits, and professional development opportunities.