Назад
Company hidden
3 дня назад

Head of Platform Product Reliability (AI Infrastructure)

Формат работы
onsite
Тип работы
fulltime
Грейд
head
Английский
b2
Страна
US
Релокация
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Head of Platform Product Reliability (AI Infrastructure): Leading system-level reliability engineering for AI servers, accelerator platforms, rack systems, and datacenter infrastructure with an accent on qualification standards, accelerated stress testing, failure analysis, and reliability modeling. Focus on building fleet reliability infrastructure, driving cross-functional corrective actions, and establishing reliability signoff criteria for large-scale deployments.

Location: On-site in San Jose, California, United States. Relocation support is available for candidates moving to San Jose.

Company

hirify.global builds inference-focused frontier AI hardware by co-designing chips, racks, software, and manufacturing systems.

What you will do

  • Define end-to-end reliability strategy for AI servers, accelerator platforms, rack systems, and datacenter infrastructure from architecture through fleet deployment.
  • Establish qualification standards, validation methodologies, EVT/DVT/PVT gates, accelerated life and stress testing, and environmental testing programs.
  • Lead root-cause investigations and corrective actions across hardware, firmware, thermal, mechanical, manufacturing, and field-deployment domains.
  • Develop MTBF projections, FIT rate analysis, Weibull lifetime models, component derating methods, and reliability growth tracking.
  • Build fleet telemetry pipelines, field feedback loops, and monitoring frameworks for deployed system health.
  • Lead reliability signoff and release-readiness reviews while building and managing the product reliability engineering organization.

Requirements

  • BS, MS, or PhD in Electrical Engineering, Mechanical Engineering, Reliability Engineering, or a related technical field.
  • 10+ years of reliability engineering experience in hardware-centric organizations, including complex system-level work.
  • Experience leading reliability programs for AI accelerators, GPU-class compute systems, hyperscale or cloud servers, networking, storage, or rack-scale infrastructure.
  • Deep understanding of thermal, power delivery, mechanical, connector, and interconnect failure mechanisms.
  • Hands-on experience with FMEA, Weibull analysis, HALT/HASS, qualification planning, failure analysis, reliability statistics, and modeling.
  • Strong technical judgment and communication skills for influencing engineering, program, and executive stakeholders.

Nice to have

  • Experience with liquid-cooled systems, high-density power delivery, or thermal management for high-power AI infrastructure.
  • Experience supporting hyperscale or cloud datacenter deployments, customer-facing reliability commitments, and SLA management.
  • Experience building reliability organizations, processes, tooling, and team norms from an early stage.
  • Experience with fleet telemetry, large-scale field reliability analytics, ODM/JDM partnerships in Taiwan or Asia, or GPU and high-speed digital systems.

Culture & Benefits

  • Fully in-person work in San Jose with close collaboration across engineering, research, manufacturing, supply chain, and datacenter operations.
  • Medical, dental, and vision coverage, plus a $500 monthly credit for waiving medical benefits.
  • $2,000 monthly housing subsidy for employees living within walking distance of the office.
  • Relocation support for moving to San Jose, wellness benefits, and daily lunch and dinner at the office.
  • Unlimited compute budget subject to ROI justification.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →