Назад
1 день назад

Hardware Failure Analysis Engineer (AI Infrastructure)

Формат работы
onsite
Тип работы
fulltime
Грейд
junior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Hardware Failure Analysis Engineer (AI Infrastructure): Analyze firmware, hardware specifications, and failures supporting AI data center operations with an accent on diagnostics, reliability validation, and vendor RMA management. Focus on proving intermittent hardware defects, developing anomaly-monitoring tools, and responding to hardware incidents in the Memphis data center.

Location: Memphis, Tennessee, or Southaven, Mississippi, United States

Company

xAI develops AI systems designed to understand the universe and support humanity's pursuit of knowledge.

What you will do

  • Analyze firmware packages and hardware specifications for compatibility, performance, reliability, security vulnerabilities, and safety risks before deployment.
  • Investigate hardware failures, including ambiguous and intermittent grey failures, using rigorous testing and data analysis.
  • Manage vendor relationships, RMA claims, negotiations, and resolution tracking.
  • Collaborate with data center operations technicians to troubleshoot, repair, and optimize servers, GPUs, networking equipment, and related systems.
  • Build monitoring tools, scripts, and processes that detect hardware anomalies and reduce downtime.
  • Document failure modes, root-cause analyses, reliability models, RMA outcomes, and hardware evaluations.

Requirements

  • Bachelor's degree in Systems Engineering, Electrical Engineering, Computer Science, or a related field, or equivalent experience.
  • At least 2 years of experience in hardware reliability engineering, preferably in high-performance computing or data center environments.
  • Expertise in firmware analysis, hardware specification review, and release validation.
  • Experience with RMA processes, vendor negotiations, and complex hardware failure diagnosis.
  • Proficiency in Python and Bash scripting, plus general experience with C, C++, Java, Rust, or a similar systems language.
  • Ability to work collaboratively with operations technicians and participate in on-call rotations and hardware incident response.

Nice to have

  • Experience with AI/ML infrastructure or supercomputing environments.
  • Knowledge of NVIDIA, Dell, HP, or Supermicro vendor ecosystems and supply chain management.
  • Hardware engineering or reliability certifications such as CRE or CompTIA Server+.
  • Experience in a fast-paced startup or technology company.

Culture & Benefits

  • Work in a small, highly motivated organization focused on engineering excellence.
  • Contribute directly through a hands-on role and a flat organizational structure.
  • Collaborate across functions with strong emphasis on concise, accurate communication.
  • Participate in on-call rotations supporting data center operations.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →