Назад
4 дня назад

Hardware Diagnostics Engineer (AI)

Формат работы
onsite
Тип работы
fulltime
Грейд
middle
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Hardware Diagnostics Engineer (AI): Running burn-in and stress testing for GPU servers, triaging hardware failures, managing out-of-band server operations, and driving vendor RMAs with an accent on reliable AI compute infrastructure. Focus on isolating faulty components, analyzing firmware and telemetry patterns, automating repetitive diagnostics, and maintaining production readiness across the fleet.

Location: Las Vegas, Nevada, United States; on-site

Company

TensorWave delivers a secure, reliable, and resilient cloud platform for AI compute at scale.

What you will do

  • Run server and GPU burn-in, stress testing, and production-readiness validation.
  • Triage failures across GPUs, memory, drives, NICs, power supplies, cabling, and other server components.
  • Manage servers out-of-band through IPMI and Redfish, including power control, boot configuration, BIOS settings, and event-log collection.
  • Apply qualified firmware baselines, maintain asset and replacement records in NetBox, and identify recurring fleet failure patterns.
  • Drive vendor RMAs from initial ticket through replacement, installation, and return of failed parts.
  • Improve runbooks, automate repetitive diagnostic tasks, and support datacenter turn-ups, expansions, and hardware escalations.

Requirements

  • 3–6 years of experience in datacenter operations, systems administration, hardware support, or infrastructure engineering.
  • Hands-on experience with enterprise server hardware, component replacement, POST and boot failures, and rack-level troubleshooting.
  • Practical experience with BMCs and out-of-band management, including IPMI, Redfish, iDRAC, iLO, or equivalent.
  • Strong Linux troubleshooting skills covering boot processes, drivers, devices, and diagnostic tools such as dmesg, lspci, ipmitool, and SMART.
  • Working scripting ability in Bash or Python, experience with hardware RMAs, and clear written communication for tickets, runbooks, and vendor cases.
  • Employment requires authorization to work in the United States.

Nice to have

  • GPU server experience, particularly with AMD GPUs and ROCm.
  • Experience with burn-in, stress testing, or node-validation tooling in GPU or HPC environments.
  • Familiarity with staged firmware updates, NetBox or other DCIM/IPAM tools, Ansible, or Python REST API automation.
  • Experience in a high-volume hardware environment such as a hyperscaler, colocation provider, systems integrator, or manufacturing test operation.

Culture & Benefits

  • Stock options.
  • 100% paid medical, dental, and vision insurance for employees.
  • Company HSA contributions, flexible spending account, disability insurance, life insurance, and supplemental insurance options.
  • 401(k), flexible PTO, paid holidays, parental leave, and employee assistance program.
  • In-office perks and participation in an on-call and hardware escalation rotation.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →