Назад
1 месяц назад

Principal Technical Program Manager (TPM) - AI Infrastructure Operations

Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Principal Technical Program Manager (TPM) - AI Infrastructure Operations (AI/HPC): Driving complex operational programs for high-scale AI and HPC infrastructure, including GPU fleet growth, data center build-outs, firmware rollouts, and Infiniband network operations with an accent on availability, uptime, and infrastructure readiness. Focus on coordinating hardware, platform, network, SRE, data center, and vendor teams, improving incident and change management, and reducing MTTR through metrics-driven execution.

Location: Houston, New York, San Francisco, or Seattle, United States

Company

Nscale is a fast-growing technology startup building AI infrastructure for large-scale customer workloads.

What you will do

  • Lead strategic operational programs covering AI infrastructure build-outs, GPU fleet software and firmware rollouts, and operational tooling with SRE.
  • Define and track infrastructure KPIs, including 97.5% availability and 99% uptime targets, with dashboards and leadership reporting.
  • Optimize incident management, change management, and postmortem workflows across Fleet Operations, Network Operations, and SRE.
  • Coordinate Hardware, Compute Platform, Network, Data Center Operations, and external GPU and networking vendors.
  • Translate capacity planning into infrastructure delivery and readiness roadmaps, ensuring GPUs, NICs, and switches meet go-live criteria.
  • Identify technical, schedule, and resource risks and develop mitigation strategies for infrastructure scaling and stability.

Requirements

  • 5+ years of experience in Technical Program Management for large-scale infrastructure or software engineering programs.
  • Strong understanding of data center infrastructure, distributed systems, Linux, and networking.
  • Experience with Agile or Scrum program management methodologies; PMP certification is preferred.
  • Experience defining and improving operational metrics such as uptime, availability, MTTR, SLOs, and SLIs.
  • Strong organizational, communication, and presentation skills, with the ability to manage ambiguity and multiple priorities.

Nice to have

  • Experience with data center infrastructure build-outs and hardware commissioning.
  • Knowledge of AI/HPC infrastructure, NVIDIA GPUs, InfiniBand/RDMA networks, and tightly coupled systems.
  • Experience supporting 24/7 mission-critical services in a hyperscale or public cloud environment.
  • Familiarity with SRE principles, infrastructure automation, and CI/CD pipelines.
  • Bachelor’s or Master’s degree in Computer Science, Engineering, or a related technical field.

Culture & Benefits

  • Collaborative, supportive, and innovative working environment.
  • Competitive package including base salary and equity.
  • Compensation reviews every 12 months.
  • Progression plan tailored to individual ambitions, with opportunities to lead and take ownership.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →