Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Head of GPU Capacity Planning (AI): Owning end-to-end capacity planning for AMD Instinct GPU clusters across U.S. data center sites with an accent on fleet inventory, demand forecasting, sellable capacity, and cross-functional allocation decisions. Focus on building auditable capacity models, anticipating infrastructure constraints, and aligning GPU deployments, sales commitments, supply chain, and long-range expansion.
Location: Remote with monthly travel to Vegas; candidates must have authorization to work in the United States.
Company
TensorWave provides a secure, reliable cloud platform for AI compute at scale, operating AMD Instinct GPU clusters across U.S. data center sites.
What you will do
- Own the authoritative view of total, committed, reserved, and available-to-sell GPU capacity across sites and GPU generations.
- Reconcile physical fleet inventory with logical and contracted allocations, including installed equipment, burn-in units, spares, and RMA inventory.
- Build rolling demand, contracted-growth, allocation, and multi-year capacity forecasts, identifying shortfalls and required lead times.
- Partner with Sales and Deal Desk to validate deal feasibility, prevent oversubscription, and publish current sellable capacity.
- Coordinate with Infrastructure Operations, Global Operations, supply chain, Finance, PMO, and data center teams on deployments, reservations, power, cooling, networking, and site expansion.
- Run the weekly capacity review, maintain auditable allocation records, and define requirements for capacity, inventory, and allocation tooling.
Requirements
- 5+ years in capacity planning, supply and demand planning, S&OP, technical program management, or infrastructure operations.
- Working knowledge of data center and compute infrastructure, including racks, power and cooling constraints, servers, GPUs, and networking.
- Advanced modeling and forecasting skills, with strong data fluency; SQL and BI tools are advantageous.
- Experience translating technical constraints into commercial guidance while working across Sales, Finance, and Operations.
- Excellent written communication and the ability to run recurring operating cadences and produce trusted reporting.
- Authorization to work in the United States is required.
Nice to have
- Experience in GPU cloud, HPC, hyperscale, colocation, or semiconductor capacity environments.
- Familiarity with Slurm, Kubernetes cluster capacity, GPU fleet management, and utilization metrics.
- Experience standing up DCIM, capacity, or inventory tooling.
- PMP or equivalent program management certification.
Culture & Benefits
- Stock options.
- 100% employer-paid medical, dental, and vision insurance for employees.
- Health Savings Account contributions, Flexible Spending Account, and supplemental health benefits.
- Paid short- and long-term disability insurance, life insurance options, and Employee Assistance Program.
- Flexible PTO, paid holidays, parental leave, and in-office perks.
- Equal opportunity workplace with reasonable accommodations available during the hiring process.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
12 дней назад
Senior Systems Engineer (AI)
150 000 - 190 000$
Lambda
14 дней назад
Data Center Manager (AI)
137 000 - 183 000$
12 дней назад
Senior Systems Engineer (GPU / HPC Infrastructure)
150 000$
11 дней назад
VP of Systems Engineering (AI Infrastructure) - West Coast
11 дней назад
Network Engineer (AI)
202 000 - 261 000$
13 дней назад