Назад
1 месяц назад

Site Reliability Engineer - Datacenter

Тип работы
fulltime
Грейд
middle
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Site Reliability Engineer - Datacenter (Hardware Reliability/Firmware): Analyzing firmware and hardware specifications, diagnosing datacenter hardware failures, and improving reliability for AI infrastructure with an accent on release validation, vulnerability analysis, vendor RMAs, and hardware diagnostics. Focus on proving intermittent failures, developing anomaly monitoring, documenting reliability models, and responding to hardware incidents in the Memphis datacenter.

Location: Memphis, Tennessee or Southaven, Mississippi, United States

Company

xAI develops AI systems designed to understand the universe and support humanity’s pursuit of knowledge.

What you will do

  • Analyze firmware packages and hardware specifications for compatibility, performance, reliability, security vulnerabilities, and safety risks.
  • Investigate hardware failures, including intermittent and “grey failures,” using rigorous testing, data analysis, logic analyzers, and diagnostic software.
  • Manage vendor relationships, RMA claims, negotiations, and resolution tracking.
  • Collaborate with datacenter operations technicians to troubleshoot, repair, and optimize servers, GPUs, and networking equipment.
  • Build monitoring tools, scripts, and processes to detect hardware anomalies and reduce downtime.
  • Document failure modes, root-cause analyses, reliability models, RMA outcomes, and hardware evaluations; participate in hardware on-call and incident response.

Requirements

  • Bachelor’s degree in Systems Engineering, Electrical Engineering, Computer Science, or a related field, or equivalent experience.
  • 2+ years of experience in hardware reliability engineering, preferably in high-performance computing or datacenter environments.
  • Expertise in firmware analysis, hardware specification review, and release validation.
  • Experience with RMA processes, vendor negotiations, and complex hardware failure diagnosis.
  • Familiarity with datacenter hardware and proficiency in Python or Bash scripting, plus experience with a systems language such as C, C++, Java, or Rust.
  • Ability to collaborate with cross-functional operations teams and apply a data-driven approach to reliability engineering.

Nice to have

  • Experience with AI/ML infrastructure or supercomputing environments.
  • Knowledge of NVIDIA, Dell, HP, or Supermicro vendor ecosystems and supply chain management.
  • Hardware engineering or reliability certifications such as CRE or CompTIA Server+.
  • Experience in a fast-paced startup or technology company.

Culture & Benefits

  • Small, highly motivated organization focused on engineering excellence.
  • Flat structure with hands-on contribution expected from all employees.
  • Leadership opportunities based on initiative and consistent delivery.
  • Work emphasizes curiosity, strong prioritization, communication, and direct contribution to the company mission.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →