Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Hardware Failure Analysis Engineer (AI Infrastructure): Analyze firmware, hardware specifications, and failures supporting AI data center operations with an accent on diagnostics, reliability validation, and vendor RMA management. Focus on proving intermittent hardware defects, developing anomaly-monitoring tools, and responding to hardware incidents in the Memphis data center.
Location: Memphis, Tennessee, or Southaven, Mississippi, United States
Company
xAI develops AI systems designed to understand the universe and support humanity's pursuit of knowledge.
What you will do
- Analyze firmware packages and hardware specifications for compatibility, performance, reliability, security vulnerabilities, and safety risks before deployment.
- Investigate hardware failures, including ambiguous and intermittent grey failures, using rigorous testing and data analysis.
- Manage vendor relationships, RMA claims, negotiations, and resolution tracking.
- Collaborate with data center operations technicians to troubleshoot, repair, and optimize servers, GPUs, networking equipment, and related systems.
- Build monitoring tools, scripts, and processes that detect hardware anomalies and reduce downtime.
- Document failure modes, root-cause analyses, reliability models, RMA outcomes, and hardware evaluations.
Requirements
- Bachelor's degree in Systems Engineering, Electrical Engineering, Computer Science, or a related field, or equivalent experience.
- At least 2 years of experience in hardware reliability engineering, preferably in high-performance computing or data center environments.
- Expertise in firmware analysis, hardware specification review, and release validation.
- Experience with RMA processes, vendor negotiations, and complex hardware failure diagnosis.
- Proficiency in Python and Bash scripting, plus general experience with C, C++, Java, Rust, or a similar systems language.
- Ability to work collaboratively with operations technicians and participate in on-call rotations and hardware incident response.
Nice to have
- Experience with AI/ML infrastructure or supercomputing environments.
- Knowledge of NVIDIA, Dell, HP, or Supermicro vendor ecosystems and supply chain management.
- Hardware engineering or reliability certifications such as CRE or CompTIA Server+.
- Experience in a fast-paced startup or technology company.
Culture & Benefits
- Work in a small, highly motivated organization focused on engineering excellence.
- Contribute directly through a hands-on role and a flat organizational structure.
- Collaborate across functions with strong emphasis on concise, accurate communication.
- Participate in on-call rotations supporting data center operations.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
4 дня назад
Forward Deployed Engineer (AI Infrastructure)
4 дня назад
RMA and Repair Manager (AI Hardware)
5 дней назад
Staff Engineer, Architecture & Performance Research Engineer for Data Center and Agentic AI CPU (RISC-V)
163 000 - 253 000$
11 часов назад
Software Engineer (AI)
5 дней назад
Design Verification Engineer (ASIC)
4 дня назад
Principal Power Design Engineer (AI Hardware)
160 000 - 240 000$