Назад
Company hidden
1 час назад

Network Operations Engineer (AI)

173 000 - 279 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
middle/senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Network Operations Engineer (AI): Managing and automating large-scale network infrastructure for AI compute with an accent on fault diagnostics, RMA pipelines, and fleet-wide tooling. Focus on building self-healing network systems and maintaining high-availability monitoring for hyperscale data centers.

Location

Must be based in San Francisco, CA; Austin, TX; New York, NY; or Seattle, WA (On-site role)

Salary: $173,000 – $279,000 per year plus equity

Company

hirify.global is building civilization-scale infrastructure for AI, focused on delivering massive compute capacity through rapid data center design and operation.

What you will do

  • Carry the on-call pager for the network fleet and manage end-to-end repair processes.
  • Develop Python and Go tooling to automate link diagnostics and remote command execution.
  • Build and maintain automated repair pipelines, including RMA initiation and ticket integration.
  • Maintain real-time monitoring and alerting platforms for hyperscale infrastructure.
  • Validate new sites and hardware through qualification testing before production traffic.

Requirements

  • Must be authorized to work in the United States.
  • Proven experience carrying a pager for a production network and managing outages independently.
  • Proficiency in writing Python or Go scripts to automate repetitive network tasks.
  • Hands-on experience with link diagnostics, optics, and protocols such as gNMI, gRPC, NETCONF, and SONiC.
  • Comfortable working at the CLI on switches and routers across a large fleet.
  • Ability to quickly reach competence in unfamiliar parts of the stack and document findings.

Nice to have

  • Experience with RMA and repair lifecycle automation.
  • Knowledge of large-scale datacenter fabric (BGP, ECMP, spine-leaf).
  • Familiarity with out-of-band network management.
  • Fluency with AI coding tools like Claude Code or Cursor.

Culture & Benefits

  • Full-time role with competitive base salary and equity.
  • Opportunity to work on high-intensity, civilization-scale AI infrastructure problems.
  • Culture of extreme ownership, velocity, and first-principles thinking.
  • Comprehensive benefits package included.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →