3 часа назад
Staff Network Production Operations Engineer
195 000 - 235 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Staff Network Production Operations Engineer (AI Infrastructure): Owning production reliability across edge, backbone, data center fabric, and GPU cluster networks supporting large-scale AI workloads with an accent on incident response, observability, and operational automation. Focus on designing remediation workflows, solving complex network reliability issues, and operating RDMA/RoCE fabrics across multi-region device fleets.
Location: On-site in San Francisco, CA; Bellevue, WA; or Sunnyvale, CA, US
Salary: $195,000–$235,000 per year plus bonus and restricted stock units
Company
builds vertically integrated energy and AI infrastructure, operating systems across power, data centers, cloud services, and AI computing.
What you will do
- Own production reliability across global edge, backbone, data center fabric, and GPU cluster networks.
- Lead and contribute to high-severity incident response, mitigation, stakeholder communication, and postmortems.
- Drive root cause analyses, identify systemic issues, and track remediation plans through completion.
- Improve network observability using streaming telemetry, SNMP, NetFlow, Kentik, Grafana, Prometheus, and ThousandEyes.
- Author runbooks, escalation playbooks, and standard operating procedures for network operations.
- Build Python-based automation, contribute to SLI/SLO definition, and mentor senior engineers.
Requirements
- 8+ years of production network engineering experience focused on operations, incident response, and reliability in large-scale or internet-scale environments.
- Hands-on experience with streaming telemetry, SNMP, NetFlow/sFlow, Grafana, Prometheus, and ThousandEyes.
- Experience operating RDMA/RoCE lossless fabrics for GPU or HPC workloads, including PFC, ECN, and DCQCN tuning.
- Expert knowledge of BGP, EVPN-VXLAN, IS-IS, OSPF, MPLS, QoS, and TCP/IP in production data center environments.
- Proficiency with Arista EOS and Juniper Junos in leaf-spine CLOS architectures and multi-vendor environments.
- Python proficiency, large device fleet operations, on-call work, and a bachelor's degree in a related field or equivalent practical experience.
Nice to have
- Experience with NVIDIA/Mellanox networking platforms in GPU cluster environments.
- Familiarity with Kentik or Arbor for traffic analysis and DDoS visibility.
- Experience defining SLIs and SLOs with SRE or product teams.
- Experience operating 10K+ device fleets or contributing to organization-wide post-incident learning programs.
Culture & Benefits
- Competitive compensation, equity, and restricted stock units.
- Paid time off, holidays, leave programs, and parental leave.
- Health, dental, vision, HSA contributions, life insurance, and disability coverage.
- Professional development, tuition reimbursement, and mental health support.
- 401(k) plan with company matching up to 4% of salary.
- Commuter benefits, cell phone stipend, meals allowance, volunteer time off, and global travel insurance.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
1 день назад
WAN Network Engineer
Datadog
4 часа назад
Staff Engineer - Cloud Networks
244 000 - 305 000$
1 день назад
IT Engineer
150 000 - 250 000$
6 дней назад
IT Operations Engineer (Networking)
120 000 - 165 000$
1 день назад
Linux Device Management Engineer
160 000 - 200 000$
6 дней назад
IT Operations Engineer (Cloud/Network)
115 000 - 155 000$