Назад
15 часов назад

L3 Support Engineer (Data Center Infrastructure)

125 000 - 180 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify RU Global, списка компаний с восточно-европейскими корнями
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
L3 Support Engineer (Data Center Infrastructure) (AI/Data Center Infrastructure): Maintaining production hardware reliability across large-scale, mission-critical data center environments with an accent on server hardware, firmware, and fleet stability. Focus on leading root cause investigations, coordinating vendor remediation, validating fleet-wide fixes, and improving hardware observability and reliability metrics.

Location: Onsite at the Birmingham, Alabama data center, United States. Applicants must be authorized to work in the country where they apply.

Salary: $125,000–$180,000 per year plus an annual performance-based bonus. A base compensation range of $147,200–$183,900 USD is also listed.

Company

Nebius builds a full-stack AI cloud platform covering data and model training through production deployment, with infrastructure spanning compute, storage, networking, and applied AI.

What you will do

  • Lead root cause analysis for complex hardware and firmware failures across production fleets.
  • Act as the senior escalation point for incidents affecting system availability or performance.
  • Coordinate with vendors on diagnostics, RMAs, firmware fixes, and corrective actions.
  • Partner with engineering and onsite operations teams to validate fixes and prevent recurrence.
  • Perform hardware and firmware validation before fleet-wide rollouts.
  • Improve hardware observability, failure tracking, reporting, and long-term fleet reliability.

Requirements

  • Strong hands-on expertise with server hardware in data center or large-scale production environments.
  • Proven experience analyzing hardware and firmware failures and identifying root causes.
  • Deep understanding of CPU, memory, storage, networking, power, BMC, and related failure modes.
  • Experience working with hardware vendors, engineering teams, and onsite operations teams.
  • Strong analytical, structured problem-solving, incident management, and communication skills.
  • Ability to manage multiple concurrent investigations with production impact.

Nice to have

  • Experience with GPU-dense, AI, or high-performance computing environments.
  • Exposure to firmware lifecycle management and large-scale rollout validation.
  • Familiarity with Linux-based production systems and infrastructure tooling.
  • Experience improving fleet-wide hardware reliability metrics at scale.

Culture & Benefits

  • Medical, dental, and vision coverage.
  • 401(k) plan with company contribution.
  • Flexible paid time off and paid parental leave.
  • Professional development and learning support.
  • Collaborative international environment focused on impactful AI infrastructure.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →