5 дней назад
Senior Technical Program Manager (HPC)
142 800 - 274 800$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Technical Program Manager (HPC): Driving the operational health, reliability, and readiness of large-scale HPC infrastructure supporting Microsoft AI workloads with an accent on cluster health, incident management, service-level objectives, and cross-functional execution. Focus on coordinating critical incident response, identifying systemic reliability issues, and building durable improvements across compute, networking, storage, capacity, datacenter, and Azure teams.
Location: Mountain View, United States
Salary: USD $142,800–$274,800 per year; a separate range of USD $188,000–$304,200 applies to the San Francisco Bay Area and New York City metropolitan area.
Company
Microsoft AI builds frontier models and products supported by large-scale AI infrastructure.
What you will do
- Own programs for HPC production health, operational readiness, cluster availability, reliability, and service performance.
- Establish operating mechanisms covering cluster health, SLA/SLO reviews, NIS/RIS tracking, incident trends, risks, dependencies, and corrective actions.
- Coordinate critical incident triage, escalation, mitigation, stakeholder communication, root-cause analysis, and follow-through.
- Identify systemic reliability issues and drive durable engineering and operational improvements.
- Build cross-functional partnerships across Azure compute, networking, storage, capacity, datacenter, and platform teams.
- Define dashboards and metrics, manage production risks, and communicate health, incidents, trade-offs, and action plans to leadership.
Requirements
- Significant experience in technical program management, infrastructure engineering, production operations, site reliability, or large-scale distributed systems.
- Experience delivering complex cross-functional programs across multiple engineering organizations.
- Technical fluency in HPC, compute, accelerators, networking, storage, schedulers, capacity, telemetry, or distributed systems.
- Experience managing production incidents, including escalation, mitigation, root-cause analysis, and corrective actions.
- Experience with SLAs, SLOs, availability metrics, and production-health indicators.
- Strong relationship-building, influence, written communication, and verbal communication skills.
Nice to have
- Experience supporting GPU or accelerator-based HPC/AI infrastructure at scale.
- Experience with cloud infrastructure and datacenter operations.
- Experience establishing reliability programs, operational scorecards, health dashboards, or incident-management mechanisms.
Culture & Benefits
- Work with HPC engineering, platform, networking, storage, capacity, datacenter, and Azure teams.
- Benefits and additional compensation may be available.
- Applications are accepted on an ongoing basis until the position is filled.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
Nscale
8 дней назад
Principal Technical Program Manager, AI Physical Deployment
250 000 - 303 000$
9 дней назад
Technical Program Manager (AI Infrastructure)
230 000 - 280 000$
7 дней назад
Program Manager (AI)
160 000 - 220 000$
6 дней назад
Lead Technical Program Manager, AI Platform
176 400 - 264 600$
8 дней назад
Principal Technical Program Manager - Hyperscale and AI Rack Systems
148 125 - 210 000$
6 дней назад
Senior Technical Program Manager (AI)
140 000 - 170 000$