3 дня назад
Sr. TPM (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Sr. TPM (AI) (AI infrastructure): Owning technical programs for data center and site operations supporting Cerebras AI Cloud and customer deployments with an accent on operational readiness, cross-functional coordination, and reliability metrics. Focus on deploying and scaling wafer-scale AI systems, leading incident postmortems, and building executive dashboards for availability, capacity, and operational risk.
Location: Sunnyvale headquarters office, on-site
Company
Systems develops wafer-scale AI hardware and infrastructure for high-speed model training and inference.
What you will do
- Own end-to-end technical programs for data center and site operations supporting AI Cloud and customer deployments.
- Coordinate Hardware and Systems Engineering, AI Cloud Infrastructure and Operations, Network and Storage Engineering, Facilities, power and cooling teams, and colocation partners.
- Drive site readiness, installation, commissioning, change management, and break/fix workflows for wafer-scale systems.
- Lead incident reviews and postmortems, ensuring corrective actions are completed.
- Define operational metrics and KPIs covering availability, reliability, incidents, MTTR, MTTD, deployment readiness, capacity, and operational risk.
- Build executive dashboards, establish governance and RACI clarity, and present status and risks to senior leadership.
Requirements
- 8+ years of experience in Technical Program Management, Infrastructure Operations, or Data Center Operations.
- Experience leading large, cross-functional infrastructure programs.
- Strong understanding of data center power and cooling, network and storage fundamentals, and hardware-centric platforms.
- Ability to define and operationalize metrics.
- Strong written and executive-level communication skills.
Nice to have
- Experience with AI/ML, HPC, or accelerator-based infrastructure.
- Experience with high-density or liquid-cooled data centers.
- Experience working with colocation providers and facilities teams.
- Background in incident management, reliability, or service operations.
Culture & Benefits
- Opportunity to build an AI platform beyond traditional GPU constraints.
- Access to cutting-edge AI research, publishing, and open-source work.
- Work on a high-performance AI supercomputer platform.
- Job stability combined with startup vitality.
- Non-corporate culture focused on individual beliefs, learning, growth, and inclusion.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
3 дня назад
Senior Staff TPM (AI Infrastructure)
230 000 - 280 000$
3 дня назад
Technical Program Manager, Infrastructure (AI)
3 дня назад
Staff TPM (AI Infrastructure)
200 000 - 240 000$
3 дня назад
Staff TPM for Managed Intelligence (AI)
200 000 - 240 000$
2 дня назад
Senior Technical Program Manager (AI Infrastructure)
160 000 - 220 000$
5 дней назад
Confluent Staff Technical Program Manager (TPM)
161 000 - 299 000$