Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Principal Frontend Network Engineer (AI Infrastructure) (Ethernet/AI infrastructure): Designing and operating high-performance front-end networks for large-scale AI GPU clusters with an accent on reliability, scalability, routing, congestion management, and storage connectivity. Focus on evolving leaf-spine architectures, solving complex cross-layer incidents, and improving observability, capacity efficiency, and uptime across distributed infrastructure.
Location: Houston, New York, San Francisco, or Seattle, United States
Company
Nscale provides GPU cloud infrastructure for AI startups and enterprise customers, supporting high-performance AI development and deployment.
What you will do
- Own the technical direction and operational strategy for front-end AI infrastructure networks.
- Design and evolve large-scale Ethernet leaf-spine and Clos fabric architectures using Arista and Nokia platforms.
- Serve as the senior escalation point for complex network incidents and lead systemic fixes.
- Improve fabric reliability, performance predictability, observability, capacity efficiency, and incident reduction.
- Define standards for routing, congestion management, firmware lifecycle, automation, configuration, and change safety across Nvidia Cumulus, Arista EOS, and Nokia platforms.
- Partner with SRE, compute, storage, and network architecture teams while mentoring senior and principal network engineers.
Requirements
- 12+ years of network engineering experience focused on hyperscale data centre, cloud, or AI infrastructure networking.
- Expertise in large-scale Ethernet data centre fabrics, including leaf-spine and Clos topologies.
- Production experience with Nvidia Cumulus, Arista EOS/Etherlink, and/or Nokia 7220 IXR, 7250 IXR, or 7750 SR platforms.
- Strong knowledge of BGP, OSPF, ECMP, and EVPN-VXLAN control planes.
- Experience with long-haul circuits, DCI, optical transport, storage networking, and shared storage connectivity.
- Ability to resolve cross-layer issues and lead complex technical initiatives across teams without direct authority.
Nice to have
- Experience operating network platforms at hyperscale or large AI infrastructure scale.
- Familiarity with front-end network patterns for inference, management, and storage tiers in massive AI clusters.
- Automation and tooling experience with Python, Ansible, validation frameworks, or telemetry pipelines.
- Experience influencing platform or infrastructure strategy at significant scale.
Culture & Benefits
- Collaborative, supportive, and innovative working environment.
- Competitive package including base salary and equity, with reviews every 12 months.
- Progression plan with opportunities to lead, challenge existing approaches, and own impact.
- Potential eligibility for bonus, equity, and/or commission programs.
- Medical, dental, and vision benefits, flexible paid time off, parental leave, and retirement plan participation may be available.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →