1 день назад
Principal Network Engineer (AI Infrastructure)
270 000 - 330 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Principal Network Engineer (AI Infrastructure): Designing and operating large-scale InfiniBand, RoCE, and Ethernet network fabrics for AI training and inference workloads with an accent on low-latency architecture, automation, security, and observability. Focus on defining reference architectures, solving complex performance and stability issues, and leading network engineering standards across data center sites.
Location: Seattle, United States
Salary: $270,000–$330,000 USD per year, plus potential bonus, equity, and/or commission.
Company
Nscale provides GPU cloud infrastructure for AI startups and enterprise customers, including the platforms and networking systems that support large-scale AI workloads.
What you will do
- Define, design, validate, and evolve InfiniBand, RoCE, and high-performance Ethernet fabrics at rack, row, and data center scale.
- Set technical direction for BGP, EVPN-VXLAN, LACP, QoS, Clos, and spine-leaf architectures across multiple sites.
- Design perimeter and security infrastructure, including firewalls, NAT, VPN, security policies, high availability, and multi-tenant segmentation.
- Lead GitOps-based network automation using Python, Ansible, infrastructure-as-code, and CI/CD pipelines.
- Drive observability, telemetry, monitoring, alerting, root-cause analysis, SLOs, and operational improvements.
- Partner with deployment, data center operations, platform, systems, storage, and vendor teams while mentoring engineers and leading technical escalations.
Requirements
- 10+ years of network engineering experience, including significant experience in HPC, AI, hyperscale, or large-scale data center environments.
- Hands-on RDMA-aware networking experience for AI/HPC workloads, including InfiniBand and/or RoCE, subnet managers such as OpenSM or UFM, and fabric orchestration.
- Expert knowledge of BGP, EVPN-VXLAN, Clos/spine-leaf architectures, and production platforms such as Cumulus, Nokia, or Arista EOS.
- Strong experience with Python, Ansible, Git-based workflows, Terraform, GitLab CI, or GitHub Actions.
- Deep experience with Juniper SRX and/or Palo Alto firewalls, security policy architecture, high availability, and multi-tenant segmentation.
- Proven ability to define architecture and technical strategy, lead complex incidents, communicate trade-offs, and mentor engineers across networking, systems, storage, and AI/HPC teams.
Culture & Benefits
- Engineering culture focused on innovation, ownership, accountability, transparency, and operational excellence.
- Medical, dental, and vision benefits.
- Flexible paid time off and parental leave.
- Retirement plan participation.
- Potential bonus, equity, and/or commission eligibility.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
6 дней назад
Senior Network Engineer (AI Infrastructure)
150 000 - 190 000$
6 дней назад
Senior Network Engineer (InfiniBand / UFM)
170 000 - 210 000$
4 часа назад
Network Engineer
250 000 - 320 000$
5 дней назад
Staff Network Engineer
100 000 - 230 000$
5 часов назад
Network Engineer (AI Infrastructure)
2 дня назад
Network Engineer (AI)
202 000 - 261 000$