6 дней назад
Senior AI Engineer (Kubernetes & Customised Scheduler)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior AI Engineer (Kubernetes & Customised Scheduler) (AI infrastructure/Kubernetes): Building and operating a proprietary Kubernetes-native scheduler for massive-scale AI factories with an accent on topology-aware GPU placement, workload orchestration, and resource-aware scheduling. Focus on designing controllers, scheduling policies, observability, multi-tenancy, and production reliability across GPU, network, storage, power, and thermal constraints.
Location: Singapore or Australia (Launceston, Hobart, Sydney, Melbourne)
Company
Technologies develops and operates sustainable AI infrastructure and the AI Cloud GPU platform across Asia Pacific.
What you will do
- Design, build, operate, and continuously improve a proprietary Kubernetes-native scheduler for large-scale AI workloads.
- Develop custom resource definitions, controllers, admission webhooks, scheduling plugins, APIs, CLI tools, automation, and workload templates.
- Implement topology-aware and AI-factory-resource-aware placement using GPU, NVLink/NVSwitch, NUMA, RDMA, storage, capacity, health, power, thermal, and maintenance signals.
- Support queueing, priorities, quotas, fair sharing, reservations, preemption, gang scheduling, backfilling, retries, checkpoint-aware scheduling, and failure recovery.
- Build observability, scheduler explanations, benchmarking, testing, CI/CD, GitOps, secure multi-tenancy, and production operations.
- Collaborate with AI, inference, platform, infrastructure, networking, storage, security, operations, and product teams on the Model-to-Grid roadmap.
Requirements
- At least 5 years of experience in DevOps, SRE, platform engineering, distributed systems, cloud infrastructure, or comparable software and systems engineering roles.
- Deep practical expertise in Kubernetes architecture and operations, including control planes, scheduling, controllers, operators, CRDs, admission controllers, APIs, networking, storage, security, and multi-tenancy.
- Experience building or integrating workload schedulers, resource managers, or orchestration systems such as Kubernetes Scheduler Framework, Kueue, KAI, Volcano, YuniKorn, Slurm, Slinky, Run:ai, or proprietary systems.
- Strong Go programming skills and Python proficiency for automation, integration, tooling, and performance analysis.
- Experience with GPU-accelerated AI/ML workloads, distributed training, inference, benchmarking, GPU allocation, topology-aware placement, NVLink, NVSwitch, RDMA, RoCEv2, NUMA, and high-performance networking.
- Experience with observability, CI/CD, GitOps, infrastructure as code, cloud-native security, troubleshooting, production reliability, and incident response.
Culture & Benefits
- Work with founders and specialists in AI infrastructure, energy systems, and next-generation compute.
- Founder-led environment with accessible leadership, fast decisions, and limited bureaucracy.
- Opportunity to contribute to sustainable AI Factories designed to operate as assets to the energy grid.
- Permanent full-time employment.
- Support for growth into new technical domains and broader responsibilities.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →