Назад
Company hidden
3 дня назад

Staff TPM for Managed Intelligence (AI)

200 000 - 240 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Staff TPM for Managed Intelligence (AI): Delivering multi-quarter programs for a managed LLM inference platform, including model onboarding, inference optimization, GPU readiness, and production reliability with an accent on cross-functional execution, model serving, and multi-tenant capacity planning. Focus on validating firmware, drivers, CUDA and ROCm stacks, governing model rollouts, and solving latency, throughput, cost, and SLA challenges at production scale.

Location: On-site in San Francisco or Sunnyvale, California, United States

Salary: $200,000–$240,000 annually plus bonus and restricted stock units

Company

hirify.global builds and operates sustainable AI cloud infrastructure, including GPU infrastructure, IaaS products, and managed inference services powered by clean energy.

What you will do

  • Own multi-quarter release planning, dependency governance, and executive communication for the Managed Inference platform.
  • Drive model version rollouts, inference optimization, GPU hardware SLA readiness, and multi-tenant capacity planning from kickoff through delivery.
  • Coordinate programs across Model Engineering, IaaS, Cloud Foundations, Data Center Operations, and external model providers.
  • Identify risks involving model serving, reliability, capacity constraints, and vendor timelines.
  • Build scalable TPM execution frameworks, real-time dashboards, and data-driven executive updates.
  • Lead phase-zero planning for model onboarding on new GPU generations, including firmware, driver, CUDA, ROCm, and inference-workload validation.

Requirements

  • 7+ years of experience as a Technical Program Manager in fast-paced technical environments.
  • Working knowledge of LLM inference and model serving, including batching, quantization, and latency, throughput, and cost trade-offs.
  • Experience with multi-tenant systems, isolation, quota management, and SLA enforcement.
  • Familiarity with fine-tuning and alignment workflows for coordinating timelines and technical risks.
  • Ability to build execution models in low-structure environments and influence engineering, product, and infrastructure leaders without direct authority.
  • Exceptional written and verbal communication, including clear, data-driven updates for executive stakeholders and daily use of AI tools to improve program execution.

Nice to have

  • Experience with AI inference or training platforms and model onboarding across GPU generations.
  • Experience coaching or mentoring junior TPMs.
  • Exposure to multi-site or globally distributed engineering teams.
  • Background at a Series D–F company or an AI infrastructure team within a hyperscaler.

Culture & Benefits

  • Competitive compensation, bonus, equity, and restricted stock units.
  • Paid time off, holidays, leave programs, and parental leave.
  • Health, dental, vision, life, disability, mental health, and wellness support.
  • 401(k) plan with company match up to 4% of salary and HSA contributions.
  • Professional development, tuition reimbursement, commuter benefits, cell phone stipend, meals allowance, and global travel insurance.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →