Назад
4 дня назад

Site Reliability Engineer (AI)

Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Site Reliability Engineer (AI/data center infrastructure): Designing campus-scale monitoring, incident command, and reliability guardrails across compute, network, storage, power, and cooling with an accent on observability, alert quality, and cross-discipline operations. Focus on leading SEV response, defining SLOs and error budgets, running blameless postmortems, and closing corrective actions for the Memphis/Southaven data center campus.

Location: Memphis, Tennessee; Southaven, Mississippi, United States

Company

xAI builds AI systems designed to understand the universe and support humanity’s pursuit of knowledge.

What you will do

  • Design and own campus-scale monitoring architecture, alert suppression, and signal quality.
  • Lead SEV-class technical incidents, coordinate incident bridges with the NOC, and maintain accurate timelines and severity classification.
  • Run blameless postmortems and drive corrective actions through completion.
  • Lead reliability projects across compute, network, storage, power, cooling, and facilities telemetry.
  • Build playbooks, run game days, maintain dependency maps, and improve runbook quality with the NOC.
  • Define error budgets and availability objectives, and participate in on-call rotations for the data center campus.

Requirements

  • Bachelor’s degree in Systems Engineering, Computer Science, Electrical Engineering, or a related field, or equivalent experience.
  • 5+ years of experience in site reliability, systems engineering, or large-scale production operations, preferably in high-performance computing or data center environments.
  • Experience commanding large-scale incidents and providing calm technical leadership during incident bridges.
  • Experience designing fleet- or campus-scale monitoring and observability, including alert hygiene, suppression, and signal quality.
  • Experience across at least two of compute, network, storage, power, and cooling or facilities telemetry.
  • Proficiency in Python and Bash scripting, plus general experience with at least one systems language such as C, C++, Java, Go, or Rust.

Nice to have

  • Experience with AI/ML infrastructure or supercomputing environments.
  • Hands-on experience with SLOs, SLIs, and error budgets.
  • Experience running game days, mapping dependencies, and managing closed-loop corrective action programs.
  • Familiarity with data center hardware and plant signals, including servers, GPUs, networking, power, and cooling.
  • Experience in a fast-paced startup or technology company.

Culture & Benefits

  • Small, highly motivated organization focused on engineering excellence.
  • Flat structure with hands-on contribution expected from all employees.
  • Leadership opportunities based on initiative and consistent delivery.
  • Collaborative work across NOC, data center operations, and infrastructure engineering.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →