Technical Program Manager (AI Infrastructure)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
TL;DR
Technical Program Manager (AI Infrastructure): Managing provider relationships and infrastructure execution to ensure scaled AI compute availability with an accent on provider onboarding, capacity rollouts, and incident response. Focus on streamlining provider operations, coordinating cross-functional SRE and engineering teams, and building scalable partner playbooks.
Location: Global Remote / San Francisco, CA
Company
provides early-stage startups with access to scaled AI infrastructure, building the orchestration layer for global AI compute.
What you will do
- Manage new provider and site onboarding, capacity expansions, and remediation of hardware or network issues.
- Act as incident commander during major provider-side incidents, coordinating between SREs, providers, and affected customers.
- Maintain detailed program plans including milestones, owners, dependencies, and risks.
- Hold providers accountable to contractual commitments through structured check-ins and leadership escalations.
- Coordinate internally with SRE, Engineering, and Product teams to resolve platform-side blockers.
- Develop standardized playbooks and provider-facing standards to accelerate site onboarding.
Requirements
- Several years of experience running technical programs in infrastructure (TPM) with execution ownership.
- Proven experience in incident management involving multiple organizations.
- Technical depth in GPUs, networking, and storage to communicate effectively with SREs and data-center engineers.
- Track record of aligning external partners and internal teams without formal authority.
- Strong fundamentals in planning, risk tracking, and direct communication.
- Comfort with ambiguity and building processes from the ground up.
Nice to have
- Direct experience with neocloud, colocation, or data-center providers.
- Vendor or supplier management background involving SLAs and escalation frameworks.
- Familiarity with NVIDIA data-center GPUs, InfiniBand/RoCE, Slurm, or Kubernetes.
- Experience supporting AI research labs or large-scale GPU customers.
Culture & Benefits
- Opportunity to join an early-stage company at the center of the AI infrastructure boom.
- High ownership as the first TPM hire for the solutions engineering team.
- Competitive compensation with meaningful equity.
- Comprehensive health, dental, and vision coverage for employees and dependents.
- 401(k) and unlimited PTO.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →