Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
TL;DR
Staff Reliability Engineer (AI Hardware): Defining the reliability strategy for next-generation AI computing systems with an accent on high-performance hardware, thermal management, and production quality. Focus on leading root-cause investigations, validating advanced cooling systems, and ensuring reliability standards across the NPI process.
Location: Hybrid, based out of Toronto, Canada
Company
Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance and cost efficiency with high-performance RISC-V CPUs.
What you will do
- Define and drive the reliability strategy for AI computing systems across data center and workstation products.
- Lead root-cause investigations on failures and implement fixes across engineering and the supply chain.
- Validate advanced cooling systems, including vapor chambers, heat pipes, and direct-to-chip liquid cooling.
- Collaborate with mechanical, electrical, thermal, and software teams to ensure system durability.
- Manage reliability standards with manufacturing partners throughout the NPI process.
- Mentor other engineers and lead design reviews to raise the technical bar for the team.
Requirements
- 8+ years of experience in reliability engineering, ideally in HPC, AI hardware, or data center systems.
- Strong expertise in statistical analysis: HALT, HASS, ALT, MTBF, Weibull, and FMEA.
- Hands-on experience in thermal labs and validating emerging cooling technologies.
- Ability to communicate technical risks and trade-offs clearly to leadership.
- Must be eligible to access U.S. export-controlled technology (compliance with EAR).
- Must be based in Toronto, Canada for hybrid work.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →