8 дней назад
Senior Principal Frontend Network Engineer (AI Infrastructure)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Principal Frontend Network Engineer (AI Infrastructure) (Ethernet, AI GPU clusters): Designing and operating high-performance front-end Ethernet networks for large-scale AI GPU clusters with an accent on leaf-spine/Clos fabrics, routing, long-haul connectivity, and storage networking. Focus on solving complex cross-layer incidents, improving reliability and observability, and defining scalable operational standards across Nvidia Cumulus, Arista EOS, and Nokia platforms.
Location: Houston, New York, San Francisco, or Seattle
Company
Nscale provides high-performance GPU cloud infrastructure for AI start-ups and enterprise customers.
What you will do
- Own the technical direction and operational strategy for front-end AI infrastructure networks.
- Design and evolve large-scale Ethernet leaf-spine and Clos fabrics using Arista and Nokia platforms.
- Serve as the senior escalation point for complex network incidents and lead systemic fixes.
- Improve fabric reliability, performance predictability, observability, capacity efficiency, and operational maturity.
- Define standards for routing, congestion management, firmware lifecycle, automation, and safe network changes across Nvidia Cumulus, Arista EOS, and Nokia platforms.
- Partner with SRE, compute, storage, and network architecture teams while mentoring senior and principal engineers.
Requirements
- 12+ years of network engineering experience focused on hyperscale data centre, cloud, or AI infrastructure networking.
- Expertise in large-scale Ethernet data centre fabrics, including leaf-spine and Clos topologies.
- Production experience with Nvidia Cumulus, Arista EOS, and/or Nokia platforms at scale.
- Strong knowledge of BGP, OSPF, ECMP, and EVPN-VXLAN.
- Experience with long-haul circuits, DCI, optical transport, dark fiber, carrier Ethernet, coherent optics, and ZR/ZR+.
- Background in hyperscale storage networking and debugging cross-layer issues involving hardware, optics, routing, and applications.
Nice to have
- Experience designing front-end networks for massive AI clusters, including inference, management, and storage tiers.
- Large-scale DCI, long-haul optical, or carrier network experience.
- Automation and tooling experience with Python, Ansible, validation frameworks, or telemetry pipelines.
- Experience influencing infrastructure strategy at significant scale.
Culture & Benefits
- Collaborative, supportive, and innovative work environment.
- Competitive package including base salary and equity, with reviews every 12 months.
- Progression plan supporting technical leadership, ownership, and professional growth.
- Benefits may include medical, dental, and vision coverage, flexible paid time off, parental leave, and retirement plan participation.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
xAI
9 дней назад
Network Engineer
150 000 - 250 000$
14 дней назад
Network Engineer IV (Cisco ACI)
108 253 - 158 733$
13 дней назад
Senior Networking Engineer (AI Infrastructure)
TensorWave
14 дней назад
Principal Network Engineer (AI Infrastructure)
11 дней назад
Network Engineer (AI/HPC)
94 000 - 117 000$
13 дней назад