обновлено 10 дней назад
Senior Slurm Cluster & HPC Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Slurm Cluster & HPC Engineer (AI/HPC): Designing and operating production Slurm clusters for large-scale GPU training and cloud workloads with an accent on topology-aware scheduling, multi-tenant isolation, Kubernetes integration, and cluster reliability. Focus on validating GPU fabric placement, shifting elastic capacity between Slurm and Kubernetes, automating bare-metal infrastructure, and diagnosing hardware and job failures.
Location: Singapore, SG
Company
is a technology company focused on Bitcoin mining solutions and AI cloud infrastructure, operating mining and HPC datacenters globally.
What you will do
- Design, deploy, and operate production Slurm clusters on bare metal and virtual machines, including high availability, authentication, accounting, and live upgrades.
- Implement topology-aware GPU scheduling for InfiniBand, RoCE, and NVLink fabrics, validating placement with NCCL and multi-node training tests.
- Manage multi-tenant scheduling policies, quotas, fairshare, preemption, reservations, and fail-closed authorization.
- Lead Slinky Slurm-on-Kubernetes integrations and shift GPU capacity between batch training and Kubernetes inference workloads.
- Operate containerized job runtimes, health checks, automated draining and requeueing, rack acceptance testing, and GPU reliability diagnostics.
- Deliver reproducible infrastructure with Terraform, Ansible, PXE, Redfish, and IPMI; integrate observability, accounting, billing, customer support, and technical documentation.
Requirements
- 8+ years of HPC, systems, or cloud infrastructure engineering experience, including 4+ years operating production Slurm clusters at 100+ GPU-node scale.
- Deep hands-on experience with Slurm configuration, scheduling policies, accounting, authentication, REST APIs, and live version upgrades.
- Strong knowledge of NVIDIA GPU infrastructure, DCGM, MIG, InfiniBand/RoCE, GPUDirect RDMA, NCCL, and GPU failure diagnosis.
- Production Kubernetes experience and exposure to Slurm-on-Kubernetes stacks such as Slinky, slurm-bridge, CoreWeave SUNK, or Nebius Soperator.
- Experience with bare-metal provisioning, firmware and BIOS lifecycle management, virtualized compute, Terraform, and Ansible.
- Working knowledge of parallel storage, Python, and Bash; clear written and verbal English communication is required.
Nice to have
- Go experience for platform control-plane and Slurm/Slinky REST integrations.
- Experience with Lustre, GPFS/Spectrum Scale, WEKA, VAST, or NFS.
Culture & Benefits
- Inclusive environment that values authenticity and diverse perspectives.
- Startup spirit within a fast-growing digital asset and AI infrastructure company.
- Opportunities to contribute directly to new projects and systems.
- Autonomy, personal accountability, learning opportunities, training, mentoring, and welfare benefits.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
10 дней назад
Senior Cloud Infrastructure Engineer (AWS)
10 дней назад
Sr. Cloud Infrastructure Engineer (AWS)
12 дней назад
Senior DevOps Engineer (Singapore)
12 дней назад
Senior Cloud Infrastructure Engineer - DevOps (AWS)
10 дней назад
Senior Cloud Infrastructure Engineer (AWS)
11 дней назад