Назад
Company hidden
6 дней назад

Senior AI Engineer (Kubernetes & Customised Scheduler)

Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Singapore/Australia
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior AI Engineer (Kubernetes & Customised Scheduler) (AI infrastructure/Kubernetes): Building and operating a proprietary Kubernetes-native scheduler for massive-scale AI factories with an accent on topology-aware GPU placement, workload orchestration, and resource-aware scheduling. Focus on designing controllers, scheduling policies, observability, multi-tenancy, and production reliability across GPU, network, storage, power, and thermal constraints.

Location: Singapore or Australia (Launceston, Hobart, Sydney, Melbourne)

Company

hirify.global Technologies develops and operates sustainable AI infrastructure and the hirify.global AI Cloud GPU platform across Asia Pacific.

What you will do

  • Design, build, operate, and continuously improve a proprietary Kubernetes-native scheduler for large-scale AI workloads.
  • Develop custom resource definitions, controllers, admission webhooks, scheduling plugins, APIs, CLI tools, automation, and workload templates.
  • Implement topology-aware and AI-factory-resource-aware placement using GPU, NVLink/NVSwitch, NUMA, RDMA, storage, capacity, health, power, thermal, and maintenance signals.
  • Support queueing, priorities, quotas, fair sharing, reservations, preemption, gang scheduling, backfilling, retries, checkpoint-aware scheduling, and failure recovery.
  • Build observability, scheduler explanations, benchmarking, testing, CI/CD, GitOps, secure multi-tenancy, and production operations.
  • Collaborate with AI, inference, platform, infrastructure, networking, storage, security, operations, and product teams on the Model-to-Grid roadmap.

Requirements

  • At least 5 years of experience in DevOps, SRE, platform engineering, distributed systems, cloud infrastructure, or comparable software and systems engineering roles.
  • Deep practical expertise in Kubernetes architecture and operations, including control planes, scheduling, controllers, operators, CRDs, admission controllers, APIs, networking, storage, security, and multi-tenancy.
  • Experience building or integrating workload schedulers, resource managers, or orchestration systems such as Kubernetes Scheduler Framework, Kueue, KAI, Volcano, YuniKorn, Slurm, Slinky, Run:ai, or proprietary systems.
  • Strong Go programming skills and Python proficiency for automation, integration, tooling, and performance analysis.
  • Experience with GPU-accelerated AI/ML workloads, distributed training, inference, benchmarking, GPU allocation, topology-aware placement, NVLink, NVSwitch, RDMA, RoCEv2, NUMA, and high-performance networking.
  • Experience with observability, CI/CD, GitOps, infrastructure as code, cloud-native security, troubleshooting, production reliability, and incident response.

Culture & Benefits

  • Work with founders and specialists in AI infrastructure, energy systems, and next-generation compute.
  • Founder-led environment with accessible leadership, fast decisions, and limited bureaucracy.
  • Opportunity to contribute to sustainable AI Factories designed to operate as assets to the energy grid.
  • Permanent full-time employment.
  • Support for growth into new technical domains and broader responsibilities.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →