4 дня назад
Senior HPC Infrastructure Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior HPC Infrastructure Engineer (AI): Building and validating fault-tolerant bare-metal Kubernetes and Slurm GPU clusters with an accent on provisioning, RDMA networking, and distributed AI workload performance. Focus on tuning NCCL, UCX, GPUDirect, GPU scheduling, hardware topology, and benchmarking large-scale AI infrastructure.
Location: Australia — Sydney, NSW or Launceston, TAS
Company
develops sustainable AI infrastructure and complex software-defined infrastructure platforms.
What you will do
- Design and implement bare-metal provisioning workflows using Ironic, Kubernetes CRDs, and custom operators.
- Deploy and manage GPU-enabled AI compute nodes with RDMA, InfiniBand, and RoCE networking.
- Optimise Kubernetes and Slurm platforms for multi-node AI training, including NCCL, UCX, GPUDirect, and fabric tuning.
- Build GPU scheduling, isolation, resource-management, observability, and cluster validation capabilities.
- Develop benchmarking workloads using MLPerf, NCCL tests, microbenchmarks, and throughput or latency validation.
- Collaborate with SRE, site operations, and networking teams on reliability, hardware bring-up, documentation, and continuous improvement.
Requirements
- Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field.
- Experience with bare-metal provisioning tools such as Metal3, OpenStack Ironic, MaaS, xCAT, or similar.
- Deep knowledge of Kubernetes internals, including CRDs, controllers, operators, and cluster lifecycle management.
- Strong understanding of Slurm, GPU systems, CUDA/NCCL, NVLink, NVSwitch, PCIe, and distributed AI or HPC workloads.
- Practical Linux systems engineering experience, plus automation with Ansible, Helm, Terraform/OpenTofu, or equivalent.
- Experience with firmware, BIOS, BMC/IPMI/Redfish, programming in Go, Bash, Rust, or Python, and production on-call support.
Culture & Benefits
- Full-time employment.
- Inclusive workplace welcoming candidates from diverse backgrounds.
- Work focused on sustainable engineering practices and AI infrastructure innovation.
- Reporting to the Senior Manager, Software Defined Infrastructure.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
2 дня назад
HPC Systems Engineer (AI)
136 300 - 231 700$
TensorWave
3 дня назад
AWS Cloud Engineer (AI/HPC)
4 дня назад
Senior Infrastructure Engineer (AI Cloud)
4 дня назад
Infrastructure Engineer (Robotics/AI)
200 000 - 350 000$
4 дня назад
Infrastructure Engineer (AI)
4 дня назад