Назад
2 дня назад

Member of Technical Staff, Compute Orchestration & Scheduling (AI)

119 800 - 234 700$
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Member of Technical Staff, Compute Orchestration & Scheduling (AI): Building and scaling orchestration, scheduling, resource allocation, and fault-recovery systems for Ray and Kubernetes workloads across large GPU clusters with an accent on topology-aware placement, distributed training, and accelerator utilization. Focus on scaling NVIDIA and AMD clusters beyond thousands of GPUs, accelerating environment initialization, recovering long-running AI jobs, and improving observability and reliability.

Location: Mountain View, United States; designated Microsoft office attendance at least four days per week for employees living within the applicable distance

Salary: USD $119,800–$234,700 per year across the U.S. for IC4 roles; higher ranges apply to IC5 and specific locations.

Company

Microsoft AI develops AI systems intended to advance science, education, productivity, and global well-being.

What you will do

  • Develop and tune scalable pretraining software for NVIDIA Grace Blackwell, Vera Rubin, and AMD MIxxx architectures.
  • Scale GPU clusters beyond thousands of GPUs and improve compute utilization.
  • Build orchestration, workload scheduling, resource allocation, quota management, initialization, and fault-recovery systems for Ray and Kubernetes workloads.
  • Collect data and insights to shape GPU and CPU compute roadmaps for large-scale AI labs.
  • Collaborate with researchers, model engineers, hardware architects, and infrastructure teams to deliver scalable platform capabilities.
  • Contribute to AI models powering Microsoft AI products and iterate quickly on solutions for users.

Requirements

  • Bachelor's degree in Computer Science or a related technical field and 6+ years of technical engineering experience with coding.
  • Experience with one or more of C, C++, Python, Go, or JavaScript, or equivalent experience.
  • Ability to work from the Mountain View, United States office at least four days per week when within the applicable commuting distance.
  • Strong systems engineering experience involving large-scale distributed infrastructure, scheduling, scaling, or fault tolerance.

Nice to have

  • Master's degree and 8+ years of technical engineering experience, or bachelor's degree and 12+ years of experience.
  • Expertise in Ray, Kubernetes, Kueue, Volcano, or another AI-focused scheduling, scaling, or fault-tolerance system.

Culture & Benefits

  • Fast-paced, design-driven product development environment.
  • Opportunity to work on infrastructure supporting frontier AI research and product development.
  • Close collaboration across research, model engineering, hardware, and infrastructure disciplines.
  • Benefits and additional compensation may be available depending on role eligibility.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →