Назад
Company hidden
6 часов назад

RAS Validation Engineer (GPU)

150 000 - 225 000$
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
RAS Validation Engineer (GPU): Leading end-to-end reliability, availability, and serviceability validation for GPU server platforms operating in Low Earth Orbit with an accent on fault injection, detection coverage, and autonomous recovery. Focus on analyzing failures across GPU, CPU, memory, interconnect, firmware, and operating-system layers while validating BMC, IPMI, Redfish, and MCTP/PLDM health visibility.

Location: San Carlos, California or Seattle, Washington. U.S. work authorization is required under ITAR export control regulations.

Salary: $150,000–$225,000 annually.

Company

hirify.global builds satellite infrastructure for powering on-orbit computing and transmitting energy and high-bandwidth optical data in Low Earth Orbit.

What you will do

  • Lead RAS validation strategy and execution for GPU server platforms, including fault injection, detection coverage, and recovery verification.
  • Analyze hardware failures and silicon errata with GPU system designers and silicon partners across DDR, HBM, CPU, and GPU subsystems.
  • Characterize fault propagation from hardware detection through firmware and operating-system layers.
  • Validate BMC and out-of-band hardware health visibility through IPMI, Redfish, and MCTP/PLDM.
  • Debug failures across GPU and CPU architecture, memory subsystems, PCIe/NVLink fabrics, and system management firmware.
  • Drive root-cause analysis, define RAS coverage metrics, and maintain traceability from fault models to test coverage.

Requirements

  • Bachelor's degree in Electrical Engineering or a related discipline.
  • 5+ years of experience in hardware validation, platform reliability engineering, or silicon validation for server-class compute systems.
  • Deep knowledge of CPU and GPU architecture, DDR and HBM memory, cache hierarchies, and PCIe, NVLink, or XGMI interconnects.
  • Strong understanding of RAS concepts, including ECC, fault containment, error propagation, MCA/MCI, and recovery mechanisms.
  • Hands-on experience with hardware, firmware, and software fault injection, plus BMC, IPMI, Redfish, and MCTP/PLDM.
  • Strong Python or equivalent scripting skills for test automation and log analysis, with experience working with silicon vendors or ODM partners.

Culture & Benefits

  • Equity in hirify.global.
  • Medical, dental, and vision insurance for employees and eligible dependents.
  • 401(k) retirement savings plan.
  • Paid time off, 10 paid holidays, and paid parental leave.
  • Relocation assistance when applicable.
  • Daily office lunch and a stocked kitchen with beverages and snacks.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →