Назад
Company hidden
5 дней назад

Senior Kubernetes Platform Engineer (AI)

Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Singapore/Australia
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Kubernetes Platform Engineer (AI/Kubernetes): Operating and evolving a fleet-scale, multi-tenant Kubernetes platform for GPU-powered AI infrastructure with an accent on cluster lifecycle management, control-plane recovery, tenant onboarding, and GPU integration. Focus on diagnosing internals-level failures, building guarded automation and fleet orchestration, and leading technical recovery during major incidents.

Location: Based in Australia or Singapore, with travel to Australian AI Factory sites as required

Company

AI infrastructure company operating large-scale GPU-powered AI Factory sites and an integrated AI cloud platform.

What you will do

  • Operate and continuously improve a multi-tenant Kubernetes platform deployed across the infrastructure estate.
  • Execute cluster provisioning, patching, upgrades, decommissioning, staged releases, and version compliance.
  • Recover Kubernetes control planes, including etcd state, certificates, failed upgrades, and corrupted resources.
  • Operate virtual clusters, tenant isolation patterns, automated onboarding pipelines, and GPU integration with device plugins and scheduling.
  • Diagnose scheduling, CNI, CSI, admission, and resource contention faults while improving validation, CI/CD, provisioning, and testing frameworks.
  • Lead technical recovery during major incidents, document runbooks, and drive permanent fixes through problem management and platform engineering.

Requirements

  • 8+ years of overall experience with substantial ownership of production Kubernetes platforms in 24/7 environments.
  • Deep experience with fleet-scale Kubernetes operations, multi-cluster lifecycle management, upgrades, and control-plane internals including etcd, API servers, controllers, schedulers, and credential rotation.
  • Production experience writing Kubernetes controllers, operators, or admission logic.
  • Experience with multi-tenant or virtual cluster platforms, GPU-enabled Kubernetes, device plugins, GPU scheduling, and driver coordination.
  • Strong infrastructure automation, infrastructure-as-code, GitOps, and operational programming skills using tools such as OpenTofu or Terraform, Ansible, Argo CD, Go, Python, or Bash.
  • Experience with major incident response, on-call escalation, post-incident reviews, runbooks, admission control, workload identity, and network policy.

Nice to have

  • Experience operating GPU or HPC Kubernetes workloads at scale.
  • Experience with automated tenant or customer onboarding pipelines and vendor Kubernetes distributions for accelerated computing.
  • Familiarity with DPU or SmartNIC networking for Kubernetes CNI design.
  • Open-source Kubernetes ecosystem contributions.
  • Bachelor's degree in computer science, engineering, or a related discipline, or equivalent experience and training.

Culture & Benefits

  • Hands-on engineering role within a 24/7 operations function.
  • Shared after-hours escalation roster for the Kubernetes estate.
  • Progressive delivery through peer review, automated testing, and staged or canary rollouts.
  • Work includes travel to Australian AI Factory sites as required.

Hiring process

  • Technical evaluation focused on Kubernetes internals, fleet operations, automation, incident recovery, and multi-tenant GPU platforms.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →