Назад
Company hidden
10 часов назад

Senior AI Infrastructure Engineer (Virtualisation)

Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Singapore/Australia
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior AI Infrastructure Engineer (Virtualisation) (AI infrastructure, storage, and Kubernetes): Designing, building, and operating software-defined infrastructure for large-scale AI workloads with an accent on virtualisation, bare-metal provisioning, high-performance storage, and GPU clusters. Focus on developing multi-tenant control planes, validating RDMA-enabled storage performance, and building Kubernetes operators and orchestration frameworks for reliable AI infrastructure.

Location: Singapore or Australia, including Melbourne, Sydney, or Launceston

Company

hirify.global develops software-defined infrastructure and sustainable solutions for large-scale AI workloads.

What you will do

  • Design and implement scalable, multi-tenant control planes for AI and infrastructure workloads.
  • Develop and operate exabyte-scale S3-compatible object storage, distributed file systems, and high-performance filesystems.
  • Provision and manage bare-metal infrastructure using platforms such as Base Command Manager, Warewulf, Ironic, and MaaS.
  • Work with RDMA, GPU Direct Storage, RoCE, InfiniBand, DPDK, Ceph, Weka, DAOS, Kubernetes, and composable storage clusters.
  • Monitor, debug, benchmark, and optimise internal clusters and storage platforms in collaboration with SRE, operations, and networking teams.
  • Build automation, validation, CI/CD, Kubernetes operators, and orchestration frameworks for large-scale GPU cluster commissioning.

Requirements

  • 6–10 years of experience in infrastructure engineering and/or storage engineering.
  • Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field.
  • Hands-on experience with bare-metal provisioning and software-defined storage platforms such as Ceph, Weka, Vast Data, DAOS, or Lustre.
  • Strong knowledge of Linux systems engineering, Kubernetes, cloud-native infrastructure, distributed systems, networking, and high-performance environments.
  • Experience with automation tools such as Ansible, Helm, Terraform/OpenTofu, or equivalent, and programming in Go, Bash, Rust, or Python.
  • Experience supporting production services through an on-call rotation and documenting architecture, procedures, and performance results.

Culture & Benefits

  • Full-time employment.
  • Collaboration across engineering, operations, SRE, site operations, and networking teams.
  • Focus on continuous technical improvement, knowledge transfer, and innovation in AI and HPC infrastructure.
  • Commitment to diversity and inclusion.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →