Назад
5 дней назад

Senior Technical Program Manager (HPC)

142 800 - 274 800$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Technical Program Manager (HPC): Driving the operational health, reliability, and readiness of large-scale HPC infrastructure supporting Microsoft AI workloads with an accent on cluster health, incident management, service-level objectives, and cross-functional execution. Focus on coordinating critical incident response, identifying systemic reliability issues, and building durable improvements across compute, networking, storage, capacity, datacenter, and Azure teams.

Location: Mountain View, United States

Salary: USD $142,800–$274,800 per year; a separate range of USD $188,000–$304,200 applies to the San Francisco Bay Area and New York City metropolitan area.

Company

Microsoft AI builds frontier models and products supported by large-scale AI infrastructure.

What you will do

  • Own programs for HPC production health, operational readiness, cluster availability, reliability, and service performance.
  • Establish operating mechanisms covering cluster health, SLA/SLO reviews, NIS/RIS tracking, incident trends, risks, dependencies, and corrective actions.
  • Coordinate critical incident triage, escalation, mitigation, stakeholder communication, root-cause analysis, and follow-through.
  • Identify systemic reliability issues and drive durable engineering and operational improvements.
  • Build cross-functional partnerships across Azure compute, networking, storage, capacity, datacenter, and platform teams.
  • Define dashboards and metrics, manage production risks, and communicate health, incidents, trade-offs, and action plans to leadership.

Requirements

  • Significant experience in technical program management, infrastructure engineering, production operations, site reliability, or large-scale distributed systems.
  • Experience delivering complex cross-functional programs across multiple engineering organizations.
  • Technical fluency in HPC, compute, accelerators, networking, storage, schedulers, capacity, telemetry, or distributed systems.
  • Experience managing production incidents, including escalation, mitigation, root-cause analysis, and corrective actions.
  • Experience with SLAs, SLOs, availability metrics, and production-health indicators.
  • Strong relationship-building, influence, written communication, and verbal communication skills.

Nice to have

  • Experience supporting GPU or accelerator-based HPC/AI infrastructure at scale.
  • Experience with cloud infrastructure and datacenter operations.
  • Experience establishing reliability programs, operational scorecards, health dashboards, or incident-management mechanisms.

Culture & Benefits

  • Work with HPC engineering, platform, networking, storage, capacity, datacenter, and Azure teams.
  • Benefits and additional compensation may be available.
  • Applications are accepted on an ongoing basis until the position is filled.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →