Назад
Company hidden
2 дня назад

Principal AI Infrastructure Engineer, Kubernetes

Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Australia
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Principal AI Infrastructure Engineer, Kubernetes (Kubernetes/GPU infrastructure): Designing and operating production-grade Kubernetes platforms across GPU-accelerated bare-metal environments with an accent on cluster lifecycle, networking, storage, security, and observability. Focus on building operators and automation, integrating NVIDIA GPU infrastructure and high-performance networking, and solving complex multi-tenant distributed-systems failures at AI-factory scale.

Location: Australia — Sydney, NSW or Launceston, TAS

Company

hirify.global Technologies designs, builds, and operates energy-efficient AI infrastructure and the hirify.global AI Cloud GPU platform across the Asia-Pacific region.

What you will do

  • Define and own the Kubernetes reference architecture for management and workload clusters, including lifecycle management, multi-tenancy, workload isolation, and failure-domain design.
  • Build backend services, APIs, controllers, operators, and automation for provisioning, configuring, upgrading, scaling, and retiring clusters.
  • Engineer bare-metal Kubernetes deployment workflows using infrastructure-as-code and automated provisioning technologies such as Cluster API, kubeadm, Redfish, PXE, Ironic, and Metal3.
  • Design networking, ingress, service discovery, DNS, load balancing, network policy, service mesh, persistent storage, backup, restore, and disaster recovery capabilities.
  • Productionise NVIDIA GPU infrastructure, accelerator scheduling, telemetry, topology-aware placement, and high-performance networking for distributed AI workloads.
  • Set engineering and operational standards, define observability and service-level objectives, lead complex incident diagnosis, and mentor senior engineers.

Requirements

  • 10+ years of infrastructure, systems, or platform engineering experience, including substantial ownership of production Kubernetes platforms and at least 3 years at senior staff, principal, or equivalent level.
  • Deep knowledge of Kubernetes internals, highly available multi-cluster platforms, bare-metal or hybrid infrastructure, and control-plane failure modes.
  • Strong Go software engineering skills with practical Python and Bash experience, including Kubernetes operators, controllers, admission webhooks, CLIs, or platform services.
  • Expert Linux systems knowledge and strong Kubernetes networking, security, governance, infrastructure automation, and GitOps experience.
  • Experience with GPU-enabled Kubernetes, NVIDIA GPU Operator, RDMA networking, distributed storage such as Ceph, and production observability.
  • CKA-level expertise, a relevant degree or equivalent practical experience, and clear technical communication across platform, networking, security, and operations teams.

Culture & Benefits

  • Full-time employment with reporting to the Head of AI Platform.
  • Work on sustainable AI infrastructure and energy-efficient GPU computing.
  • Inclusive workplace encouraging applications from candidates of all backgrounds.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →