Назад
Company hidden
4 дня назад

Senior HPC Infrastructure Engineer (AI)

Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Australia
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior HPC Infrastructure Engineer (AI): Building and validating fault-tolerant bare-metal Kubernetes and Slurm GPU clusters with an accent on provisioning, RDMA networking, and distributed AI workload performance. Focus on tuning NCCL, UCX, GPUDirect, GPU scheduling, hardware topology, and benchmarking large-scale AI infrastructure.

Location: Australia — Sydney, NSW or Launceston, TAS

Company

hirify.global develops sustainable AI infrastructure and complex software-defined infrastructure platforms.

What you will do

  • Design and implement bare-metal provisioning workflows using Ironic, Kubernetes CRDs, and custom operators.
  • Deploy and manage GPU-enabled AI compute nodes with RDMA, InfiniBand, and RoCE networking.
  • Optimise Kubernetes and Slurm platforms for multi-node AI training, including NCCL, UCX, GPUDirect, and fabric tuning.
  • Build GPU scheduling, isolation, resource-management, observability, and cluster validation capabilities.
  • Develop benchmarking workloads using MLPerf, NCCL tests, microbenchmarks, and throughput or latency validation.
  • Collaborate with SRE, site operations, and networking teams on reliability, hardware bring-up, documentation, and continuous improvement.

Requirements

  • Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field.
  • Experience with bare-metal provisioning tools such as Metal3, OpenStack Ironic, MaaS, xCAT, or similar.
  • Deep knowledge of Kubernetes internals, including CRDs, controllers, operators, and cluster lifecycle management.
  • Strong understanding of Slurm, GPU systems, CUDA/NCCL, NVLink, NVSwitch, PCIe, and distributed AI or HPC workloads.
  • Practical Linux systems engineering experience, plus automation with Ansible, Helm, Terraform/OpenTofu, or equivalent.
  • Experience with firmware, BIOS, BMC/IPMI/Redfish, programming in Go, Bash, Rust, or Python, and production on-call support.

Culture & Benefits

  • Full-time employment.
  • Inclusive workplace welcoming candidates from diverse backgrounds.
  • Work focused on sustainable engineering practices and AI infrastructure innovation.
  • Reporting to the Senior Manager, Software Defined Infrastructure.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →