6 дней назад
Senior Network Production Operations Engineer (AI Infrastructure)
165 000 - 200 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Network Production Operations Engineer (AI Infrastructure): Supporting production reliability across edge, backbone, data center, and GPU cluster networks with an accent on incident response, Python automation, observability, and SLI/SLO execution. Focus on operating RDMA/RoCE fabrics, diagnosing high-severity network events, and maintaining reliable connectivity for large-scale AI workloads.
Location: On-site in San Francisco, CA; Bellevue, WA; or Sunnyvale, CA, US
Salary: $165,000–$200,000 annually, plus bonus and restricted stock units
Company
builds vertically integrated energy and AI infrastructure, operating systems from energy generation through cloud services to support large-scale AI workloads.
What you will do
- Support production reliability across global edge, backbone, data center fabric, and GPU cluster networks.
- Build and maintain Python tooling to automate remediation, diagnostics, and operational workflows.
- Respond to high-severity network incidents, including detection, triage, mitigation, stakeholder communication, and postmortems.
- Perform root cause analysis and drive network remediation items to completion.
- Develop observability automation using streaming telemetry, SNMP, NetFlow, Kentik, Grafana, Prometheus, and ThousandEyes.
- Maintain runbooks, escalation playbooks, SOPs, dashboards, alerts, and reliability metrics with Architecture and SRE teams.
Requirements
- 5+ years of production network engineering experience focused on operations, incident response, and reliability in large-scale environments.
- Strong Python and scripting skills for diagnostic tooling and automation.
- Experience with streaming telemetry, SNMP, NetFlow/sFlow, Grafana, Prometheus, ThousandEyes, and related APIs.
- Experience operating RDMA/RoCE lossless fabrics for GPU or HPC workloads, including PFC, ECN, and DCQCN tuning.
- Production knowledge of BGP, EVPN-VXLAN, IS-IS, OSPF, MPLS, QoS, TCP/IP, Arista EOS, and Junos in leaf-spine CLOS architectures.
- Bachelor’s degree in Computer Science, Electrical Engineering, or a related field, or equivalent practical experience; willingness to support on-call operations.
Nice to have
- Experience with NVIDIA/Mellanox networking platforms in GPU clusters.
- Familiarity with Kentik or Arbor for traffic analysis and DDoS visibility.
- Experience contributing to SLIs and SLOs with SRE or product teams.
- Experience operating fleets of 10,000+ devices in hyperscale or cloud environments.
Culture & Benefits
- Competitive compensation with bonus, restricted stock units, and a 401(k) plan with employer matching up to 4% of salary.
- Health, dental, and vision insurance, including employer HSA contributions.
- Paid time off, paid holidays, parental leave, volunteer time off, and paid life and disability insurance.
- Professional development, tuition reimbursement, and mental health and wellness support.
- Commuter benefits for parking and transit, plus a cell phone stipend.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
13 дней назад
Site Reliability Engineer (AI)
200 000 - 400 000$
Baseten
8 дней назад
Site Reliability Engineer (AI)
165 000 - 330 000$
12 дней назад
Senior Cloud Infrastructure and Networking
125 000 - 135 000$
11 дней назад
Staff Triage & RCA Engineer (Robotics)
191 000 - 249 000$
10 дней назад
Staff Site Reliability Engineer (Cybersecurity)
199 750 - 270 000$
8 дней назад
Production Engineer (Cybersecurity)
102 400 - 128 000$