Назад
5 часов назад

Principal Frontend Network Engineer (AI Infrastructure)

Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Principal Frontend Network Engineer (AI Infrastructure) (Ethernet/AI infrastructure): Designing and operating high-performance front-end networks for large-scale AI GPU clusters with an accent on reliability, scalability, routing, congestion management, and storage connectivity. Focus on evolving leaf-spine architectures, solving complex cross-layer incidents, and improving observability, capacity efficiency, and uptime across distributed infrastructure.

Location: Houston, New York, San Francisco, or Seattle, United States

Company

Nscale provides GPU cloud infrastructure for AI startups and enterprise customers, supporting high-performance AI development and deployment.

What you will do

  • Own the technical direction and operational strategy for front-end AI infrastructure networks.
  • Design and evolve large-scale Ethernet leaf-spine and Clos fabric architectures using Arista and Nokia platforms.
  • Serve as the senior escalation point for complex network incidents and lead systemic fixes.
  • Improve fabric reliability, performance predictability, observability, capacity efficiency, and incident reduction.
  • Define standards for routing, congestion management, firmware lifecycle, automation, configuration, and change safety across Nvidia Cumulus, Arista EOS, and Nokia platforms.
  • Partner with SRE, compute, storage, and network architecture teams while mentoring senior and principal network engineers.

Requirements

  • 12+ years of network engineering experience focused on hyperscale data centre, cloud, or AI infrastructure networking.
  • Expertise in large-scale Ethernet data centre fabrics, including leaf-spine and Clos topologies.
  • Production experience with Nvidia Cumulus, Arista EOS/Etherlink, and/or Nokia 7220 IXR, 7250 IXR, or 7750 SR platforms.
  • Strong knowledge of BGP, OSPF, ECMP, and EVPN-VXLAN control planes.
  • Experience with long-haul circuits, DCI, optical transport, storage networking, and shared storage connectivity.
  • Ability to resolve cross-layer issues and lead complex technical initiatives across teams without direct authority.

Nice to have

  • Experience operating network platforms at hyperscale or large AI infrastructure scale.
  • Familiarity with front-end network patterns for inference, management, and storage tiers in massive AI clusters.
  • Automation and tooling experience with Python, Ansible, validation frameworks, or telemetry pipelines.
  • Experience influencing platform or infrastructure strategy at significant scale.

Culture & Benefits

  • Collaborative, supportive, and innovative working environment.
  • Competitive package including base salary and equity, with reviews every 12 months.
  • Progression plan with opportunities to lead, challenge existing approaches, and own impact.
  • Potential eligibility for bonus, equity, and/or commission programs.
  • Medical, dental, and vision benefits, flexible paid time off, parental leave, and retirement plan participation may be available.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →