Назад
14 часов назад

Principal Network Engineer (AI Infrastructure)

Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
UK
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Principal Network Engineer (AI Infrastructure): Designing and operating large-scale InfiniBand, RoCE, and Ethernet networks for AI training and inference workloads with an accent on high-performance fabric architecture, automation, security, and observability. Focus on defining reference architectures, solving complex reliability and performance issues, and setting technical standards across multi-site data centre infrastructure.

Location: UK

Company

Nscale provides GPU cloud infrastructure for AI start-ups and enterprise customers, supporting high-performance AI development and workloads at scale.

What you will do

  • Define, design, validate, and evolve InfiniBand, RoCE, and Ethernet fabric architectures across racks, rows, and data centres.
  • Set technical direction for BGP, EVPN-VXLAN, LACP, QoS, and high-performance network integration with bare-metal provisioning and cluster management.
  • Lead GitOps-based network automation using Python, Ansible, Infrastructure-as-Code, and CI/CD workflows.
  • Design perimeter security, firewalls, NAT, VPNs, security policies, and multi-tenant network segmentation.
  • Establish SLOs, observability, telemetry, monitoring, alerting, runbooks, and operational standards for network services.
  • Lead architecture reviews, complex incidents, technical escalations, cross-functional decisions, and engineer mentoring.

Requirements

  • 10+ years of network engineering experience in HPC, AI, hyperscale, or large-scale data centre environments.
  • Hands-on RDMA-aware networking experience with InfiniBand and/or RoCE, including OpenSM or NVIDIA UFM.
  • Expert knowledge of BGP, EVPN-VXLAN, Clos/spine-leaf architectures, and production platforms such as Cumulus, Nokia, or Arista EOS.
  • Strong Python and Ansible automation skills, plus Git workflows and Infrastructure-as-Code or CI/CD tools such as Terraform, GitLab CI, or GitHub Actions.
  • Deep experience designing firewall infrastructure with Juniper SRX and/or Palo Alto, as well as telemetry and observability for high-throughput environments.
  • Ability to define architecture and technical strategy, lead complex incidents, influence without formal authority, communicate trade-offs, and mentor engineers.

Culture & Benefits

  • Collaborative, supportive, and innovation-focused engineering environment.
  • Highly competitive base salary and equity package with reviews every 12 months.
  • Flexible workplace with autonomy over daily scheduling.
  • Progression plan aligned with individual ambitions and opportunities to influence a global AI platform.
  • Inclusive environment supporting applications from underrepresented groups and workplace accommodations.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →