Назад
обновлено 4 дня назад

Staff Reliability Engineer (AI Hardware)

Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Canada
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Staff Reliability Engineer (AI Hardware): Defining reliability strategy for next-generation AI computing systems with an accent on MTBF modeling, accelerated life testing, and cross-functional hardware validation. Focus on leading root-cause investigations, overseeing HALT/HASS testing, assessing design risks, and converting test findings into improvements for high-performance AI platforms.

Location: Hybrid, based in Toronto, Canada; travel to third-party test facilities and manufacturing partners, including sites in Taiwan, is required.

Company

Develops high-performance AI computing platforms combining software, compilers, networking, CPUs, and semiconductor technologies.

What you will do

  • Define the reliability strategy for next-generation AI computing systems.
  • Develop MTBF models, predictive reliability frameworks, and accelerated life testing plans.
  • Lead root-cause investigations and drive corrective actions across engineering and the supply chain.
  • Partner with hardware, software, mechanical, electrical, thermal, Systems Engineering, and Compliance Validation teams.
  • Plan and oversee HALT/HASS testing at third-party facilities, resolve device-under-test issues, and assess design risks.
  • Feed test findings into system design improvements and support validation and certification readiness.

Requirements

  • 8+ years of experience in reliability engineering, preferably in high-performance computing, AI hardware, or data center systems.
  • Experience with statistical reliability methods, HALT, HASS, ALT, MTBF, Weibull analysis, and FMEA.
  • Ability to investigate technical problems in thermal lab environments and communicate risks and trade-offs to leadership.
  • Bachelor’s or Master’s degree in Mechanical Engineering, Electrical Engineering, Reliability Engineering, or a related field.
  • Ability to work in a hybrid role based in Toronto, Canada, with travel to testing and manufacturing sites.

Culture & Benefits

  • Work on high-performance AI computing systems and a RISC-V CPU platform.
  • Collaborative environment spanning hardware, software, manufacturing, and systems engineering.
  • Highly competitive compensation package and benefits.
  • Equal opportunity employment.
  • Employment is contingent upon eligibility to access U.S. export-controlled technology.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →