Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Product Manager, Network (AI): Owns product strategy and delivery for operational software running a global GPU fleet, with an accent on provisioning, fleet health, incident response, networking, and hardware lifecycle management. Focus on designing infrastructure products, improving availability and time-to-recover, and coordinating multi-quarter initiatives across Fleet Software, SRE, data centre operations, and Support.
Location: New York, United States
Salary: $220,000–$260,000 USD per year, with potential bonus, equity, and/or commission.
Company
Nscale is building a vertically integrated GenAI cloud platform spanning data centres, software, and applications for the AI stack.
What you will do
- Own the strategy and roadmap for a major Fleet Operations product area, including provisioning, fleet health, incident and repair workflows, firmware, lifecycle management, capacity, or inventory.
- Lead cross-functional initiatives from problem framing through rollout across live GPU clusters.
- Translate operational issues and recurring support work into tooling, automation, and platform capabilities.
- Define and manage fleet metrics such as availability, utilisation, MTTR, time-to-bring-up, hardware failure rates, and ticket deflection.
- Partner with engineering on architecture and trade-offs across bare metal, orchestration, observability, control planes, and network fabrics.
- Drive incident reviews and postmortems into product commitments, mentor product managers, and represent Fleet Operations in planning and leadership reviews.
Requirements
- 5–8 years of product management experience in software or technology, with ownership of significant infrastructure, platform, or operations-facing products.
- Strong technical fluency in large-scale systems, including provisioning, orchestration, observability, and control-plane design.
- Experience building products for SREs, NOC or support teams, data centre technicians, or similar operators.
- Experience with data centre networking technologies, including InfiniBand and RoCE, backend and frontend network fabrics, WAN, edge, backbone architectures, peering, and traffic engineering.
- Ability to turn ambiguous operational problems into shipped products that improve reliability, efficiency, or time-to-recover, with strong written and verbal communication skills.
- Experience partnering with network engineering teams on fabric health, congestion monitoring, and link-level failure workflows.
Nice to have
- Degree in computer science, engineering, or a related field, or prior experience as an engineer or SRE.
- Experience with cloud infrastructure, bare-metal provisioning, fleet or hardware lifecycle management, observability platforms, or incident management tooling.
- Experience with OpenStack Ironic, MAAS, Tinkerbell, NetBox, Device42, Nautobot, Jira Service Management, ServiceNow, Zendesk, Freshservice, Grafana, Prometheus, or Datadog.
- Familiarity with GPU or accelerated compute environments, data centre operations, and hyperscaler-style fleet management.
- Experience in high-growth or early-stage environments where the product is built alongside the fleet.
Culture & Benefits
- Culture focused on innovation, ownership, accountability, transparency, collaboration, adaptability, and resilience.
- Medical, dental, and vision benefits.
- Flexible paid time off and parental leave.
- Retirement plan participation and potential bonus, equity, and/or commission programs.
- Inclusive and equitable workplace with accommodation support available.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
3 дня назад
Senior Product Manager (AI Infrastructure)
200 000 - 250 000$
5 дней назад
Director of Product Management, Enterprise Core Platform (AI)
273 600 - 342 000$
Baseten
13 часов назад
Product Manager (AI Infrastructure)
235 000 - 335 000$
Roblox
2 дня назад
Senior Product Manager, Inference Platform (AI)
280 540 - 330 950$
Roblox
2 дня назад
Senior Product Manager, Compute Platform (AI)
280 540 - 330 950$
6 дней назад
Staff Product Manager (Network Observability)
153 400 - 231 220$