Назад
Company hidden
1 день назад

Staff Engineer, Memory Systems Architecture (DRAM/HBM)

163 000 - 253 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Staff Engineer, Memory Systems Architecture (DRAM/HBM): Developing in-field telemetry and platform RAS solutions to identify, analyze, and mitigate DRAM failures across data-center memory fleets with an accent on ECC, failure logging, and virtual root-cause analysis. Focus on designing memory fault-management algorithms, standardizing DRAM/HBM failure telemetry, and validating page offlining and hPPR solutions on real servers and applications.

Location: Daily onsite presence at the San Jose headquarters in the United States

Base pay range: $163,000–$253,000 USD per year, plus incentive opportunities.

Company

hirify.global develops memory and technology solutions for smartphones, electric vehicles, hyperscale data centers, IoT devices, and other applications.

What you will do

  • Analyze large-scale in-field memory telemetry to identify DRAM failure modes, abnormalities, and failure-rate trends.
  • Recommend solutions to mitigate field DRAM failures using knowledge of SoC controllers, memory operation, and platform RAS.
  • Communicate improved ECC schemes to customers based on Samsung DRAM failure modes, including DQ and burst failures.
  • Work with customers to establish the value of in-field fault-management architecture.
  • Contribute to OCP standardization of DRAM and HBM failure logging.
  • Design and validate platform RAS algorithms, including page offlining and hPPR, through proof-of-concept testing on real servers and applications.

Requirements

  • 10+ years of relevant industry experience with a bachelor's degree, 8+ years with a master's degree, or 5+ years with a PhD; hardware fault management, reliability, or data-center fleet management experience is required.
  • Knowledge of platform memory subsystems and RAS technologies, including ECC, page offlining, hPPR, and hardware sparing.
  • Experience with ECC design and verification, reverse engineering, and CPU-to-memory address mapping.
  • Experience modifying memory-controller registers and contributing commits to the Linux kernel.
  • Understanding of DRAM and HBM failure modes.
  • Ability to collaborate inclusively, learn independently, and develop innovative solutions.

Culture & Benefits

  • Work in an incubation team focused on improving customer quality experience and minimizing downtime in AI/ML hardware deployments.
  • Conduct research independently and collaboratively, with opportunities to publish findings through whitepapers and conferences.
  • Flexible work environment aligned with the onsite San Jose work policy.
  • Medical, dental, vision, 401(k), paid time off of 4+ weeks annually, holidays, and sick leave.
  • Family-care, emotional-wellness, charitable-giving, fitness, onsite café, and gym benefits.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →