2 дня назад
Member of Technical Staff – Capacity & Efficiency Infrastructure (AI)
119 800 - 234 700$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Member of Technical Staff – Capacity & Efficiency Infrastructure (AI) (Python/C++/GPU infrastructure): Building and optimizing distributed training infrastructure, telemetry systems, and efficiency tools for frontier-scale AI models with an accent on GPU clusters, performance analysis, and large-scale machine learning systems. Focus on profiling and debugging compute bottlenecks, optimizing collective communication across NVLink and InfiniBand, and improving the reliability and utilization of thousands of GPUs.
Location: Mountain View, United States. Employees living within 50 miles of the designated U.S. office are expected to work from the office at least four days per week.
Salary: USD $119,800–$234,700 per year for IC4, or USD $142,800–$274,800 per year for IC5. San Francisco Bay Area and New York City ranges may be higher.
Company
Microsoft AI builds AI systems and products designed to advance science, education, productivity, and global well-being.
What you will do
- Design, implement, test, and optimize distributed training infrastructure in Python and C++ for large-scale GPU clusters.
- Build telemetry systems that measure infrastructure and model performance, utilization, and cost.
- Profile, benchmark, and debug compute, memory, networking, and storage bottlenecks.
- Develop architectural improvements, automated efficiency recommendations, and tools for fleet-wide optimization.
- Optimize collective communication libraries such as NCCL for NVLink and InfiniBand topologies.
- Partner with ML researchers, infrastructure engineers, and hardware teams on distributed training and next-generation accelerators.
Requirements
- Bachelor’s degree in Computer Science or a related technical discipline, plus 6+ years of technical engineering experience, or equivalent experience.
- Experience coding in languages including C, C++, C#, Java, JavaScript, or Python.
- Deep understanding of GPU architectures and deep learning or large language model architectures.
- Experience profiling and analyzing large-scale distributed computing systems and generative AI models.
- Experience with low-level GPU programming such as CUDA, Triton, or NCCL, and frameworks such as PyTorch or JAX.
- Experience building large-scale machine learning or generative AI infrastructure, distributed training systems, networking, or storage systems.
Nice to have
- 10+ years of engineering experience with a bachelor’s degree, or 8+ years with a master’s degree.
- Experience leading technical projects and supporting architecture decisions with data.
- Track record of contributing to high-performance computing or large-scale AI infrastructure projects.
Culture & Benefits
- Work across engineering, research, and product development on frontier-scale AI models.
- Collaborate with teams building products used by billions of people.
- Participate in a fast-moving codebase focused on large-scale training infrastructure.
- Benefits and additional compensation may be available depending on the role and location.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
2 дня назад
Member of Technical Staff, Pre-Training Infrastructure (AI)
142 800 - 274 800$
6 дней назад
AI Performance Engineer
75 000 - 100 000$
3 дня назад
AI Research Engineer
100 000 - 150 000$
4 дня назад
Staff ML Engineer (AI)
130 000 - 190 000$
3 дня назад
AI Performance Engineer
100 000 - 150 000$
4 дня назад
Applied ML Engineer (AI)
200 000 - 235 000$