Назад
Company hidden
обновлено 10 дней назад

Senior Slurm Cluster & HPC Engineer (AI)

Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Singapore
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Slurm Cluster & HPC Engineer (AI/HPC): Designing and operating production Slurm clusters for large-scale GPU training and cloud workloads with an accent on topology-aware scheduling, multi-tenant isolation, Kubernetes integration, and cluster reliability. Focus on validating GPU fabric placement, shifting elastic capacity between Slurm and Kubernetes, automating bare-metal infrastructure, and diagnosing hardware and job failures.

Location: Singapore, SG

Company

hirify.global is a technology company focused on Bitcoin mining solutions and AI cloud infrastructure, operating mining and HPC datacenters globally.

What you will do

  • Design, deploy, and operate production Slurm clusters on bare metal and virtual machines, including high availability, authentication, accounting, and live upgrades.
  • Implement topology-aware GPU scheduling for InfiniBand, RoCE, and NVLink fabrics, validating placement with NCCL and multi-node training tests.
  • Manage multi-tenant scheduling policies, quotas, fairshare, preemption, reservations, and fail-closed authorization.
  • Lead Slinky Slurm-on-Kubernetes integrations and shift GPU capacity between batch training and Kubernetes inference workloads.
  • Operate containerized job runtimes, health checks, automated draining and requeueing, rack acceptance testing, and GPU reliability diagnostics.
  • Deliver reproducible infrastructure with Terraform, Ansible, PXE, Redfish, and IPMI; integrate observability, accounting, billing, customer support, and technical documentation.

Requirements

  • 8+ years of HPC, systems, or cloud infrastructure engineering experience, including 4+ years operating production Slurm clusters at 100+ GPU-node scale.
  • Deep hands-on experience with Slurm configuration, scheduling policies, accounting, authentication, REST APIs, and live version upgrades.
  • Strong knowledge of NVIDIA GPU infrastructure, DCGM, MIG, InfiniBand/RoCE, GPUDirect RDMA, NCCL, and GPU failure diagnosis.
  • Production Kubernetes experience and exposure to Slurm-on-Kubernetes stacks such as Slinky, slurm-bridge, CoreWeave SUNK, or Nebius Soperator.
  • Experience with bare-metal provisioning, firmware and BIOS lifecycle management, virtualized compute, Terraform, and Ansible.
  • Working knowledge of parallel storage, Python, and Bash; clear written and verbal English communication is required.

Nice to have

  • Go experience for platform control-plane and Slurm/Slinky REST integrations.
  • Experience with Lustre, GPFS/Spectrum Scale, WEKA, VAST, or NFS.

Culture & Benefits

  • Inclusive environment that values authenticity and diverse perspectives.
  • Startup spirit within a fast-growing digital asset and AI infrastructure company.
  • Opportunities to contribute directly to new projects and systems.
  • Autonomy, personal accountability, learning opportunities, training, mentoring, and welfare benefits.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →