2 дня назад
Principal AI Infrastructure Engineer, Kubernetes
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Principal AI Infrastructure Engineer, Kubernetes (Kubernetes/GPU infrastructure): Designing and operating production-grade Kubernetes platforms across GPU-accelerated bare-metal environments with an accent on cluster lifecycle, networking, storage, security, and observability. Focus on building operators and automation, integrating NVIDIA GPU infrastructure and high-performance networking, and solving complex multi-tenant distributed-systems failures at AI-factory scale.
Location: Australia — Sydney, NSW or Launceston, TAS
Company
Technologies designs, builds, and operates energy-efficient AI infrastructure and the AI Cloud GPU platform across the Asia-Pacific region.
What you will do
- Define and own the Kubernetes reference architecture for management and workload clusters, including lifecycle management, multi-tenancy, workload isolation, and failure-domain design.
- Build backend services, APIs, controllers, operators, and automation for provisioning, configuring, upgrading, scaling, and retiring clusters.
- Engineer bare-metal Kubernetes deployment workflows using infrastructure-as-code and automated provisioning technologies such as Cluster API, kubeadm, Redfish, PXE, Ironic, and Metal3.
- Design networking, ingress, service discovery, DNS, load balancing, network policy, service mesh, persistent storage, backup, restore, and disaster recovery capabilities.
- Productionise NVIDIA GPU infrastructure, accelerator scheduling, telemetry, topology-aware placement, and high-performance networking for distributed AI workloads.
- Set engineering and operational standards, define observability and service-level objectives, lead complex incident diagnosis, and mentor senior engineers.
Requirements
- 10+ years of infrastructure, systems, or platform engineering experience, including substantial ownership of production Kubernetes platforms and at least 3 years at senior staff, principal, or equivalent level.
- Deep knowledge of Kubernetes internals, highly available multi-cluster platforms, bare-metal or hybrid infrastructure, and control-plane failure modes.
- Strong Go software engineering skills with practical Python and Bash experience, including Kubernetes operators, controllers, admission webhooks, CLIs, or platform services.
- Expert Linux systems knowledge and strong Kubernetes networking, security, governance, infrastructure automation, and GitOps experience.
- Experience with GPU-enabled Kubernetes, NVIDIA GPU Operator, RDMA networking, distributed storage such as Ceph, and production observability.
- CKA-level expertise, a relevant degree or equivalent practical experience, and clear technical communication across platform, networking, security, and operations teams.
Culture & Benefits
- Full-time employment with reporting to the Head of AI Platform.
- Work on sustainable AI infrastructure and energy-efficient GPU computing.
- Inclusive workplace encouraging applications from candidates of all backgrounds.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
6 дней назад
Site Reliability Engineer (Cloud Banking)
6 дней назад
Senior DevSecOps Engineer (AWS)
10 часов назад
Senior Observability & Telemetry Engineer (GPU/AI Infrastructure)
Canva
5 дней назад
Staff Software Engineer (Developer Experience)
Canva
5 дней назад
Senior Software Engineer (DevX)
Canva
5 дней назад