Назад
8 дней назад

Senior Principal Frontend Network Engineer (AI Infrastructure)

Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Principal Frontend Network Engineer (AI Infrastructure) (Ethernet, AI GPU clusters): Designing and operating high-performance front-end Ethernet networks for large-scale AI GPU clusters with an accent on leaf-spine/Clos fabrics, routing, long-haul connectivity, and storage networking. Focus on solving complex cross-layer incidents, improving reliability and observability, and defining scalable operational standards across Nvidia Cumulus, Arista EOS, and Nokia platforms.

Location: Houston, New York, San Francisco, or Seattle

Company

Nscale provides high-performance GPU cloud infrastructure for AI start-ups and enterprise customers.

What you will do

  • Own the technical direction and operational strategy for front-end AI infrastructure networks.
  • Design and evolve large-scale Ethernet leaf-spine and Clos fabrics using Arista and Nokia platforms.
  • Serve as the senior escalation point for complex network incidents and lead systemic fixes.
  • Improve fabric reliability, performance predictability, observability, capacity efficiency, and operational maturity.
  • Define standards for routing, congestion management, firmware lifecycle, automation, and safe network changes across Nvidia Cumulus, Arista EOS, and Nokia platforms.
  • Partner with SRE, compute, storage, and network architecture teams while mentoring senior and principal engineers.

Requirements

  • 12+ years of network engineering experience focused on hyperscale data centre, cloud, or AI infrastructure networking.
  • Expertise in large-scale Ethernet data centre fabrics, including leaf-spine and Clos topologies.
  • Production experience with Nvidia Cumulus, Arista EOS, and/or Nokia platforms at scale.
  • Strong knowledge of BGP, OSPF, ECMP, and EVPN-VXLAN.
  • Experience with long-haul circuits, DCI, optical transport, dark fiber, carrier Ethernet, coherent optics, and ZR/ZR+.
  • Background in hyperscale storage networking and debugging cross-layer issues involving hardware, optics, routing, and applications.

Nice to have

  • Experience designing front-end networks for massive AI clusters, including inference, management, and storage tiers.
  • Large-scale DCI, long-haul optical, or carrier network experience.
  • Automation and tooling experience with Python, Ansible, validation frameworks, or telemetry pipelines.
  • Experience influencing infrastructure strategy at significant scale.

Culture & Benefits

  • Collaborative, supportive, and innovative work environment.
  • Competitive package including base salary and equity, with reviews every 12 months.
  • Progression plan supporting technical leadership, ownership, and professional growth.
  • Benefits may include medical, dental, and vision coverage, flexible paid time off, parental leave, and retirement plan participation.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →