Назад
Company hidden
20 дней назад

Senior AI Scheduling & Orchestration Engineer

Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Singapore
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior AI Scheduling & Orchestration Engineer (Kubernetes/HPC): Designing and operating scheduling and orchestration systems for large-scale AI workloads with an accent on gang scheduling, accelerator allocation, topology-aware placement, and GPU sharing. Focus on solving resource contention and deadlock scenarios, scaling HPC scheduling stacks, and integrating orchestration with GPU hardware and storage systems.

Location: Singapore, SG

Company

hirify.global is a technology company focused on Bitcoin mining, AI cloud, ASIC manufacturing, datacenter operations, and high-performance computing infrastructure.

What you will do

  • Design and implement batch scheduling architectures with Volcano or YuniKorn for multi-node gang scheduling.
  • Develop cluster-wide admission control and job queueing with Kueue for high-volume AI workloads.
  • Use Kubernetes Dynamic Resource Allocation and custom scheduler plugins for accelerator requests.
  • Build topology-aware pod placement strategies for NVLink and InfiniBand environments.
  • Implement GPU sharing with MIG and time-slicing, alongside multi-tenancy isolation policies.
  • Improve scheduling reliability and scalability, resolve resource contention and deadlocks, and collaborate with GPU Systems and Storage teams.

Requirements

  • Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, or a related field.
  • 6+ years of distributed systems engineering experience with hands-on expertise in Kubernetes scheduling frameworks and orchestrators.
  • Experience with AI workload execution and distributed training frameworks such as PyTorch Distributed, Ray, or MPI.
  • Experience operating, debugging, and scaling scheduling stacks in HPC or large-scale production cloud environments.
  • Strong knowledge of GPU architectures and scheduling challenges for distributed AI training and inference.
  • Experience with infrastructure automation and infrastructure-as-code, including Terraform or Go-based Operators, plus technical leadership and communication skills.

Culture & Benefits

  • Inclusive environment valuing authenticity and diverse perspectives.
  • Startup-oriented culture within a fast-growing company.
  • Opportunities to contribute to new projects and digital asset industry infrastructure.
  • Autonomy, personal accountability, professional growth, training, and mentoring.
  • Welfare benefits and developmental opportunities.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →