Назад
Company hidden
3 дня назад

Sr. TPM (AI)

Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Sr. TPM (AI) (AI infrastructure): Owning technical programs for data center and site operations supporting Cerebras AI Cloud and customer deployments with an accent on operational readiness, cross-functional coordination, and reliability metrics. Focus on deploying and scaling wafer-scale AI systems, leading incident postmortems, and building executive dashboards for availability, capacity, and operational risk.

Location: Sunnyvale headquarters office, on-site

Company

hirify.global Systems develops wafer-scale AI hardware and infrastructure for high-speed model training and inference.

What you will do

  • Own end-to-end technical programs for data center and site operations supporting AI Cloud and customer deployments.
  • Coordinate Hardware and Systems Engineering, AI Cloud Infrastructure and Operations, Network and Storage Engineering, Facilities, power and cooling teams, and colocation partners.
  • Drive site readiness, installation, commissioning, change management, and break/fix workflows for wafer-scale systems.
  • Lead incident reviews and postmortems, ensuring corrective actions are completed.
  • Define operational metrics and KPIs covering availability, reliability, incidents, MTTR, MTTD, deployment readiness, capacity, and operational risk.
  • Build executive dashboards, establish governance and RACI clarity, and present status and risks to senior leadership.

Requirements

  • 8+ years of experience in Technical Program Management, Infrastructure Operations, or Data Center Operations.
  • Experience leading large, cross-functional infrastructure programs.
  • Strong understanding of data center power and cooling, network and storage fundamentals, and hardware-centric platforms.
  • Ability to define and operationalize metrics.
  • Strong written and executive-level communication skills.

Nice to have

  • Experience with AI/ML, HPC, or accelerator-based infrastructure.
  • Experience with high-density or liquid-cooled data centers.
  • Experience working with colocation providers and facilities teams.
  • Background in incident management, reliability, or service operations.

Culture & Benefits

  • Opportunity to build an AI platform beyond traditional GPU constraints.
  • Access to cutting-edge AI research, publishing, and open-source work.
  • Work on a high-performance AI supercomputer platform.
  • Job stability combined with startup vitality.
  • Non-corporate culture focused on individual beliefs, learning, growth, and inclusion.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →