5 дней назад
Senior Kubernetes Platform Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Kubernetes Platform Engineer (AI/Kubernetes): Operating and evolving a fleet-scale, multi-tenant Kubernetes platform for GPU-powered AI infrastructure with an accent on cluster lifecycle management, control-plane recovery, tenant onboarding, and GPU integration. Focus on diagnosing internals-level failures, building guarded automation and fleet orchestration, and leading technical recovery during major incidents.
Location: Based in Australia or Singapore, with travel to Australian AI Factory sites as required
Company
AI infrastructure company operating large-scale GPU-powered AI Factory sites and an integrated AI cloud platform.
What you will do
- Operate and continuously improve a multi-tenant Kubernetes platform deployed across the infrastructure estate.
- Execute cluster provisioning, patching, upgrades, decommissioning, staged releases, and version compliance.
- Recover Kubernetes control planes, including etcd state, certificates, failed upgrades, and corrupted resources.
- Operate virtual clusters, tenant isolation patterns, automated onboarding pipelines, and GPU integration with device plugins and scheduling.
- Diagnose scheduling, CNI, CSI, admission, and resource contention faults while improving validation, CI/CD, provisioning, and testing frameworks.
- Lead technical recovery during major incidents, document runbooks, and drive permanent fixes through problem management and platform engineering.
Requirements
- 8+ years of overall experience with substantial ownership of production Kubernetes platforms in 24/7 environments.
- Deep experience with fleet-scale Kubernetes operations, multi-cluster lifecycle management, upgrades, and control-plane internals including etcd, API servers, controllers, schedulers, and credential rotation.
- Production experience writing Kubernetes controllers, operators, or admission logic.
- Experience with multi-tenant or virtual cluster platforms, GPU-enabled Kubernetes, device plugins, GPU scheduling, and driver coordination.
- Strong infrastructure automation, infrastructure-as-code, GitOps, and operational programming skills using tools such as OpenTofu or Terraform, Ansible, Argo CD, Go, Python, or Bash.
- Experience with major incident response, on-call escalation, post-incident reviews, runbooks, admission control, workload identity, and network policy.
Nice to have
- Experience operating GPU or HPC Kubernetes workloads at scale.
- Experience with automated tenant or customer onboarding pipelines and vendor Kubernetes distributions for accelerated computing.
- Familiarity with DPU or SmartNIC networking for Kubernetes CNI design.
- Open-source Kubernetes ecosystem contributions.
- Bachelor's degree in computer science, engineering, or a related discipline, or equivalent experience and training.
Culture & Benefits
- Hands-on engineering role within a 24/7 operations function.
- Shared after-hours escalation roster for the Kubernetes estate.
- Progressive delivery through peer review, automated testing, and staged or canary rollouts.
- Work includes travel to Australian AI Factory sites as required.
Hiring process
- Technical evaluation focused on Kubernetes internals, fleet operations, automation, incident recovery, and multi-tenant GPU platforms.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
10 дней назад
Senior Slurm Cluster & HPC Engineer (AI)
10 дней назад
Senior Cloud Infrastructure Engineer (AWS)
10 дней назад
Sr. Cloud Infrastructure Engineer (AWS)
12 дней назад
Senior DevOps Engineer (Singapore)
12 дней назад
Software Engineer, DevOps
12 дней назад