Назад
Company hidden
6 дней назад

Senior Network Production Operations Engineer (AI Infrastructure)

165 000 - 200 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Network Production Operations Engineer (AI Infrastructure): Supporting production reliability across edge, backbone, data center, and GPU cluster networks with an accent on incident response, Python automation, observability, and SLI/SLO execution. Focus on operating RDMA/RoCE fabrics, diagnosing high-severity network events, and maintaining reliable connectivity for large-scale AI workloads.

Location: On-site in San Francisco, CA; Bellevue, WA; or Sunnyvale, CA, US

Salary: $165,000–$200,000 annually, plus bonus and restricted stock units

Company

hirify.global builds vertically integrated energy and AI infrastructure, operating systems from energy generation through cloud services to support large-scale AI workloads.

What you will do

  • Support production reliability across global edge, backbone, data center fabric, and GPU cluster networks.
  • Build and maintain Python tooling to automate remediation, diagnostics, and operational workflows.
  • Respond to high-severity network incidents, including detection, triage, mitigation, stakeholder communication, and postmortems.
  • Perform root cause analysis and drive network remediation items to completion.
  • Develop observability automation using streaming telemetry, SNMP, NetFlow, Kentik, Grafana, Prometheus, and ThousandEyes.
  • Maintain runbooks, escalation playbooks, SOPs, dashboards, alerts, and reliability metrics with Architecture and SRE teams.

Requirements

  • 5+ years of production network engineering experience focused on operations, incident response, and reliability in large-scale environments.
  • Strong Python and scripting skills for diagnostic tooling and automation.
  • Experience with streaming telemetry, SNMP, NetFlow/sFlow, Grafana, Prometheus, ThousandEyes, and related APIs.
  • Experience operating RDMA/RoCE lossless fabrics for GPU or HPC workloads, including PFC, ECN, and DCQCN tuning.
  • Production knowledge of BGP, EVPN-VXLAN, IS-IS, OSPF, MPLS, QoS, TCP/IP, Arista EOS, and Junos in leaf-spine CLOS architectures.
  • Bachelor’s degree in Computer Science, Electrical Engineering, or a related field, or equivalent practical experience; willingness to support on-call operations.

Nice to have

  • Experience with NVIDIA/Mellanox networking platforms in GPU clusters.
  • Familiarity with Kentik or Arbor for traffic analysis and DDoS visibility.
  • Experience contributing to SLIs and SLOs with SRE or product teams.
  • Experience operating fleets of 10,000+ devices in hyperscale or cloud environments.

Culture & Benefits

  • Competitive compensation with bonus, restricted stock units, and a 401(k) plan with employer matching up to 4% of salary.
  • Health, dental, and vision insurance, including employer HSA contributions.
  • Paid time off, paid holidays, parental leave, volunteer time off, and paid life and disability insurance.
  • Professional development, tuition reimbursement, and mental health and wellness support.
  • Commuter benefits for parking and transit, plus a cell phone stipend.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →