Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Principal Network Engineer (AI Infrastructure): Designing and operating large-scale InfiniBand, RoCE, and Ethernet networks for AI training and inference workloads with an accent on high-performance fabric architecture, automation, security, and observability. Focus on defining reference architectures, solving complex reliability and performance issues, and setting technical standards across multi-site data centre infrastructure.
Location: UK
Company
Nscale provides GPU cloud infrastructure for AI start-ups and enterprise customers, supporting high-performance AI development and workloads at scale.
What you will do
- Define, design, validate, and evolve InfiniBand, RoCE, and Ethernet fabric architectures across racks, rows, and data centres.
- Set technical direction for BGP, EVPN-VXLAN, LACP, QoS, and high-performance network integration with bare-metal provisioning and cluster management.
- Lead GitOps-based network automation using Python, Ansible, Infrastructure-as-Code, and CI/CD workflows.
- Design perimeter security, firewalls, NAT, VPNs, security policies, and multi-tenant network segmentation.
- Establish SLOs, observability, telemetry, monitoring, alerting, runbooks, and operational standards for network services.
- Lead architecture reviews, complex incidents, technical escalations, cross-functional decisions, and engineer mentoring.
Requirements
- 10+ years of network engineering experience in HPC, AI, hyperscale, or large-scale data centre environments.
- Hands-on RDMA-aware networking experience with InfiniBand and/or RoCE, including OpenSM or NVIDIA UFM.
- Expert knowledge of BGP, EVPN-VXLAN, Clos/spine-leaf architectures, and production platforms such as Cumulus, Nokia, or Arista EOS.
- Strong Python and Ansible automation skills, plus Git workflows and Infrastructure-as-Code or CI/CD tools such as Terraform, GitLab CI, or GitHub Actions.
- Deep experience designing firewall infrastructure with Juniper SRX and/or Palo Alto, as well as telemetry and observability for high-throughput environments.
- Ability to define architecture and technical strategy, lead complex incidents, influence without formal authority, communicate trade-offs, and mentor engineers.
Culture & Benefits
- Collaborative, supportive, and innovation-focused engineering environment.
- Highly competitive base salary and equity package with reviews every 12 months.
- Flexible workplace with autonomy over daily scheduling.
- Progression plan aligned with individual ambitions and opportunities to influence a global AI platform.
- Inclusive environment supporting applications from underrepresented groups and workplace accommodations.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
TensorWave
16 часов назад
Principal Network Engineer (AI Infrastructure)
7 дней назад
Network Engineer
Circle B
2 дня назад
Senior Network Engineer (GPU Cloud Infrastructure)
1 час назад
Network Reliability Engineer (Go/Python)
23 часа назад
Network Engineer (AI/HPC)
1 день назад