Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Network Operations Center Specialist (AI Infrastructure): Monitoring campus health signals and coordinating incident response across compute, network, storage, and facilities systems with an accent on SLA-driven escalation, incident communications, and operational continuity. Focus on running incident bridges, maintaining live timelines, producing major-incident reports, and driving corrective projects to closure.
Location: Southaven, Mississippi, or Memphis, Tennessee, United States. Rotating shifts include nights and weekends for continuous campus coverage.
Company
xAI develops AI systems designed to understand the universe and support humanity’s pursuit of knowledge.
What you will do
- Monitor cluster health, node availability, network health, facility trends, storage alarms, and threshold breaches from the operations console.
- Acknowledge, classify, document, verify, and escalate alerts within SLA using the NOC escalation matrix.
- Open and lead incident bridges, provide stakeholder updates on a fixed cadence, maintain live timelines, and identify ownership delays.
- Prepare first-pass incident framing and hand off in-depth analysis to SRE and Hardware Failure Analysis.
- Run structured shift handoffs, maintain durable shift logs, and preserve cross-site awareness.
- Write major-incident reports, create corrective projects in Linear, and track them through closure.
Requirements
- Experience in a 24/7 operations environment such as a NOC, SOC, dispatch, or mission control.
- Proven ability to acknowledge, classify, and escalate incidents under SLA.
- Experience running incident bridges, providing scheduled stakeholder updates, and maintaining live timelines.
- Strong written and verbal communication skills, including writing clear updates during active incidents.
- Pattern recognition across compute, network, storage, and/or facilities signals.
- Ability to work rotating shifts, including nights and weekends, for continuous campus coverage.
Nice to have
- NOC, data center operations, or campus reliability experience in high-performance computing, AI/ML infrastructure, or large-scale production environments.
- Experience writing major-incident reports and driving corrective actions to completion.
- Familiarity with Linear or similar work-tracking tools.
- Experience partnering with SRE, SiteOps, and Facilities.
- Participation in game days, tabletop exercises, or runbook improvement programs.
Culture & Benefits
- Small, highly motivated team focused on engineering excellence.
- Flat organizational structure with hands-on contribution expected.
- Leadership opportunities based on initiative and consistent delivery.
- Operational work includes maintaining and improving runbooks, escalation matrices, and communication templates.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →