2 дня назад
Member of Technical Staff, Compute Orchestration & Scheduling (AI)
119 800 - 234 700$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Member of Technical Staff, Compute Orchestration & Scheduling (AI): Building and scaling orchestration, scheduling, resource allocation, and fault-recovery systems for Ray and Kubernetes workloads across large GPU clusters with an accent on topology-aware placement, distributed training, and accelerator utilization. Focus on scaling NVIDIA and AMD clusters beyond thousands of GPUs, accelerating environment initialization, recovering long-running AI jobs, and improving observability and reliability.
Location: Mountain View, United States; designated Microsoft office attendance at least four days per week for employees living within the applicable distance
Salary: USD $119,800–$234,700 per year across the U.S. for IC4 roles; higher ranges apply to IC5 and specific locations.
Company
Microsoft AI develops AI systems intended to advance science, education, productivity, and global well-being.
What you will do
- Develop and tune scalable pretraining software for NVIDIA Grace Blackwell, Vera Rubin, and AMD MIxxx architectures.
- Scale GPU clusters beyond thousands of GPUs and improve compute utilization.
- Build orchestration, workload scheduling, resource allocation, quota management, initialization, and fault-recovery systems for Ray and Kubernetes workloads.
- Collect data and insights to shape GPU and CPU compute roadmaps for large-scale AI labs.
- Collaborate with researchers, model engineers, hardware architects, and infrastructure teams to deliver scalable platform capabilities.
- Contribute to AI models powering Microsoft AI products and iterate quickly on solutions for users.
Requirements
- Bachelor's degree in Computer Science or a related technical field and 6+ years of technical engineering experience with coding.
- Experience with one or more of C, C++, Python, Go, or JavaScript, or equivalent experience.
- Ability to work from the Mountain View, United States office at least four days per week when within the applicable commuting distance.
- Strong systems engineering experience involving large-scale distributed infrastructure, scheduling, scaling, or fault tolerance.
Nice to have
- Master's degree and 8+ years of technical engineering experience, or bachelor's degree and 12+ years of experience.
- Expertise in Ray, Kubernetes, Kueue, Volcano, or another AI-focused scheduling, scaling, or fault-tolerance system.
Culture & Benefits
- Fast-paced, design-driven product development environment.
- Opportunity to work on infrastructure supporting frontier AI research and product development.
- Close collaboration across research, model engineering, hardware, and infrastructure disciplines.
- Benefits and additional compensation may be available depending on role eligibility.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
6 дней назад
Staff Software Engineer (Kubernetes)
215 000 - 265 000$
Lambda
3 дня назад
Senior Software Engineer (AI)
255 000 - 346 000$
Lambda
3 дня назад
Staff Engineer (Managed Kubernetes)
349 000 - 465 000$
6 дней назад
Staff Software Engineer, Cloud Infrastructure (AI)
181 000 - 265 000$
Roblox
5 дней назад
Principal Software Engineer (AI)
295 250 - 345 040$
2 дня назад
Platform Engineer II (Kubernetes)
128 520 - 173 000$