11 часов назад
Staff AI Scheduling & Orchestration Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Staff AI Scheduling & Orchestration Engineer (AI): Building a high-performance scheduling fabric for a NeoCloud platform with an accent on multi-node AI workload placement, GPU utilization, and topology-aware orchestration. Focus on designing gang scheduling and admission control, optimizing NVLink and InfiniBand placement, and solving resource contention and scalability challenges in large-scale HPC environments.
Location: Remote within San Jose, California or Austin, Texas
Company
Technology company building Bitcoin mining infrastructure and AI computational and cloud infrastructure, including data centers and advanced cloud capabilities.
What you will do
- Design and implement batch scheduling architectures with Volcano or YuniKorn for multi-node gang scheduling.
- Develop cluster-wide admission control and job queueing with Kueue for high-volume AI workloads.
- Use Kubernetes Dynamic Resource Allocation and custom scheduler plugins for accelerator management.
- Architect topology-aware pod placement optimized for NVLink and InfiniBand communication.
- Implement GPU sharing with MIG and time-slicing, along with multi-tenant isolation policies.
- Collaborate with GPU Systems and Storage teams, improve reliability and scalability, and resolve contention and deadlocks in large-scale HPC environments.
Requirements
- Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, or a related field.
- 6+ years of distributed systems engineering experience with hands-on Kubernetes scheduling and orchestration expertise.
- Experience with AI workload execution and distributed training frameworks such as PyTorch Distributed, Ray, or MPI.
- Experience operating, debugging, and scaling scheduling stacks in HPC or large-scale production cloud environments.
- Strong knowledge of GPU architectures and distributed AI training and inference scheduling.
- Experience with infrastructure automation and infrastructure-as-code, including Terraform or Go-based Operators, plus technical leadership and communication skills.
Culture & Benefits
- Full-time engineering role in a high-velocity environment.
- Cross-functional collaboration with GPU Systems and Storage teams.
- Mentorship of junior engineers and participation in design reviews.
- Equal employment opportunity in accordance with applicable country, state, and local laws.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →