Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer (AI/data center infrastructure): Designing campus-scale monitoring, incident command, and reliability guardrails across compute, network, storage, power, and cooling with an accent on observability, alert quality, and cross-discipline operations. Focus on leading SEV response, defining SLOs and error budgets, running blameless postmortems, and closing corrective actions for the Memphis/Southaven data center campus.
Location: Memphis, Tennessee; Southaven, Mississippi, United States
Company
xAI builds AI systems designed to understand the universe and support humanity’s pursuit of knowledge.
What you will do
- Design and own campus-scale monitoring architecture, alert suppression, and signal quality.
- Lead SEV-class technical incidents, coordinate incident bridges with the NOC, and maintain accurate timelines and severity classification.
- Run blameless postmortems and drive corrective actions through completion.
- Lead reliability projects across compute, network, storage, power, cooling, and facilities telemetry.
- Build playbooks, run game days, maintain dependency maps, and improve runbook quality with the NOC.
- Define error budgets and availability objectives, and participate in on-call rotations for the data center campus.
Requirements
- Bachelor’s degree in Systems Engineering, Computer Science, Electrical Engineering, or a related field, or equivalent experience.
- 5+ years of experience in site reliability, systems engineering, or large-scale production operations, preferably in high-performance computing or data center environments.
- Experience commanding large-scale incidents and providing calm technical leadership during incident bridges.
- Experience designing fleet- or campus-scale monitoring and observability, including alert hygiene, suppression, and signal quality.
- Experience across at least two of compute, network, storage, power, and cooling or facilities telemetry.
- Proficiency in Python and Bash scripting, plus general experience with at least one systems language such as C, C++, Java, Go, or Rust.
Nice to have
- Experience with AI/ML infrastructure or supercomputing environments.
- Hands-on experience with SLOs, SLIs, and error budgets.
- Experience running game days, mapping dependencies, and managing closed-loop corrective action programs.
- Familiarity with data center hardware and plant signals, including servers, GPUs, networking, power, and cooling.
- Experience in a fast-paced startup or technology company.
Culture & Benefits
- Small, highly motivated organization focused on engineering excellence.
- Flat structure with hands-on contribution expected from all employees.
- Leadership opportunities based on initiative and consistent delivery.
- Collaborative work across NOC, data center operations, and infrastructure engineering.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
2 дня назад
Senior Site Reliability Engineer (AI)
191 000 - 226 000$
Anthropic
2 дня назад
Staff+ Site Reliability Engineer (Safeguards ML Infra)
320 000 - 485 000$
6 дней назад
Senior SecDevOps Engineer (AI)
165 000 - 200 000$
7 дней назад
Systems Engineer (AI Infrastructure)
140 000 - 225 000$
5 дней назад
Site Reliability Engineer (Aerospace)
6 дней назад
Infrastructure Engineer (AI)
200 000 - 300 000$