3 дня назад
Head of Platform Product Reliability (AI Infrastructure)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Head of Platform Product Reliability (AI Infrastructure): Leading system-level reliability engineering for AI servers, accelerator platforms, rack systems, and datacenter infrastructure with an accent on qualification standards, accelerated stress testing, failure analysis, and reliability modeling. Focus on building fleet reliability infrastructure, driving cross-functional corrective actions, and establishing reliability signoff criteria for large-scale deployments.
Location: On-site in San Jose, California, United States. Relocation support is available for candidates moving to San Jose.
Company
builds inference-focused frontier AI hardware by co-designing chips, racks, software, and manufacturing systems.
What you will do
- Define end-to-end reliability strategy for AI servers, accelerator platforms, rack systems, and datacenter infrastructure from architecture through fleet deployment.
- Establish qualification standards, validation methodologies, EVT/DVT/PVT gates, accelerated life and stress testing, and environmental testing programs.
- Lead root-cause investigations and corrective actions across hardware, firmware, thermal, mechanical, manufacturing, and field-deployment domains.
- Develop MTBF projections, FIT rate analysis, Weibull lifetime models, component derating methods, and reliability growth tracking.
- Build fleet telemetry pipelines, field feedback loops, and monitoring frameworks for deployed system health.
- Lead reliability signoff and release-readiness reviews while building and managing the product reliability engineering organization.
Requirements
- BS, MS, or PhD in Electrical Engineering, Mechanical Engineering, Reliability Engineering, or a related technical field.
- 10+ years of reliability engineering experience in hardware-centric organizations, including complex system-level work.
- Experience leading reliability programs for AI accelerators, GPU-class compute systems, hyperscale or cloud servers, networking, storage, or rack-scale infrastructure.
- Deep understanding of thermal, power delivery, mechanical, connector, and interconnect failure mechanisms.
- Hands-on experience with FMEA, Weibull analysis, HALT/HASS, qualification planning, failure analysis, reliability statistics, and modeling.
- Strong technical judgment and communication skills for influencing engineering, program, and executive stakeholders.
Nice to have
- Experience with liquid-cooled systems, high-density power delivery, or thermal management for high-power AI infrastructure.
- Experience supporting hyperscale or cloud datacenter deployments, customer-facing reliability commitments, and SLA management.
- Experience building reliability organizations, processes, tooling, and team norms from an early stage.
- Experience with fleet telemetry, large-scale field reliability analytics, ODM/JDM partnerships in Taiwan or Asia, or GPU and high-speed digital systems.
Culture & Benefits
- Fully in-person work in San Jose with close collaboration across engineering, research, manufacturing, supply chain, and datacenter operations.
- Medical, dental, and vision coverage, plus a $500 monthly credit for waiving medical benefits.
- $2,000 monthly housing subsidy for employees living within walking distance of the office.
- Relocation support for moving to San Jose, wellness benefits, and daily lunch and dinner at the office.
- Unlimited compute budget subject to ROI justification.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
5 дней назад
Fleet Reliability Engineer (Maritime IoT)
3 дня назад
Senior Staff Hardware Qualification Engineer (AI)
175 000 - 265 000$
3 дня назад
Principal Design Engineer (AI Infrastructure)
200 000 - 240 000$
3 дня назад
Lead Functional Safety Engineer (Robotics)
170 000 - 259 000$
9 часов назад
Head of Quality & Reliability (Aerospace)
170 000 - 235 000$
3 дня назад
Hardware Systems Engineer (AI)
250 000 - 400 000$