Назад
1 день назад

Principal Network Engineer (AI Infrastructure)

270 000 - 330 000$
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Principal Network Engineer (AI Infrastructure): Designing and operating large-scale InfiniBand, RoCE, and Ethernet network fabrics for AI training and inference workloads with an accent on low-latency architecture, automation, security, and observability. Focus on defining reference architectures, solving complex performance and stability issues, and leading network engineering standards across data center sites.

Location: Seattle, United States

Salary: $270,000–$330,000 USD per year, plus potential bonus, equity, and/or commission.

Company

Nscale provides GPU cloud infrastructure for AI startups and enterprise customers, including the platforms and networking systems that support large-scale AI workloads.

What you will do

  • Define, design, validate, and evolve InfiniBand, RoCE, and high-performance Ethernet fabrics at rack, row, and data center scale.
  • Set technical direction for BGP, EVPN-VXLAN, LACP, QoS, Clos, and spine-leaf architectures across multiple sites.
  • Design perimeter and security infrastructure, including firewalls, NAT, VPN, security policies, high availability, and multi-tenant segmentation.
  • Lead GitOps-based network automation using Python, Ansible, infrastructure-as-code, and CI/CD pipelines.
  • Drive observability, telemetry, monitoring, alerting, root-cause analysis, SLOs, and operational improvements.
  • Partner with deployment, data center operations, platform, systems, storage, and vendor teams while mentoring engineers and leading technical escalations.

Requirements

  • 10+ years of network engineering experience, including significant experience in HPC, AI, hyperscale, or large-scale data center environments.
  • Hands-on RDMA-aware networking experience for AI/HPC workloads, including InfiniBand and/or RoCE, subnet managers such as OpenSM or UFM, and fabric orchestration.
  • Expert knowledge of BGP, EVPN-VXLAN, Clos/spine-leaf architectures, and production platforms such as Cumulus, Nokia, or Arista EOS.
  • Strong experience with Python, Ansible, Git-based workflows, Terraform, GitLab CI, or GitHub Actions.
  • Deep experience with Juniper SRX and/or Palo Alto firewalls, security policy architecture, high availability, and multi-tenant segmentation.
  • Proven ability to define architecture and technical strategy, lead complex incidents, communicate trade-offs, and mentor engineers across networking, systems, storage, and AI/HPC teams.

Culture & Benefits

  • Engineering culture focused on innovation, ownership, accountability, transparency, and operational excellence.
  • Medical, dental, and vision benefits.
  • Flexible paid time off and parental leave.
  • Retirement plan participation.
  • Potential bonus, equity, and/or commission eligibility.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →