4 дня назад
Network Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Network Engineer (AI) (RDMA, BGP, EVPN/VXLAN): Designing, deploying, and operating high-performance network fabrics for a multi-tenant AI cluster with an accent on lossless GPU networking, tenant isolation, and network automation. Focus on building leaf-spine data centre fabrics, troubleshooting end-to-end performance across hosts and infrastructure, and maintaining recoverable out-of-band and corporate networks.
Location: London, England, United Kingdom. Workplace: On-site
Company
is an energy startup building an integrated energy company across solar generation, batteries, grid infrastructure, real-time power trading, AI, hardware, and distributed energy systems.
What you will do
- Design, deploy, and operate lossless RDMA-capable fabrics for GPU compute and storage traffic, including RoCEv2, InfiniBand, QoS, congestion control, and buffer tuning.
- Build and manage leaf-spine data centre fabrics with routed underlays and BGP, EVPN, and VXLAN overlays.
- Implement tenant isolation across compute, storage, and management planes, and support tenant onboarding, segmentation, bandwidth guarantees, and capacity planning.
- Automate network provisioning, configuration, and validation using Python, Ansible, NetBox, and CI-based configuration management.
- Build telemetry and observability, troubleshoot performance from optics and cabling through switches and host NICs, and operate out-of-band recovery infrastructure.
- Own the office wired and wireless network, firewalls, VPN and remote access, office-to-data-centre connectivity, and networking documentation and knowledge sharing.
Requirements
- At least 5 years of experience operating production data centre networks.
- Strong experience with BGP, overlay and encapsulation design, EVPN/VXLAN, and leaf-spine or Clos fabrics.
- Practical experience with RDMA fabrics, including lossless Ethernet with PFC, ECN, and DCQCN, or InfiniBand.
- Experience with modern data centre network operating systems, Linux networking and administration, and network automation with Python, Ansible, source-of-truth systems, and version control.
- Experience with network telemetry and monitoring, including Prometheus/Grafana, sFlow/IPFIX, or streaming telemetry.
- Experience running corporate or campus networks, including wired and wireless switching, NAC/802.1X, VPN, remote access, and clear technical communication.
Nice to have
- GPU cluster networking with NVIDIA Spectrum-X, Quantum InfiniBand, ConnectX/BlueField, UFM, SHARP, or equivalent technologies.
- Container networking, collective-communication libraries, multi-tenant VRF isolation, enterprise firewalls, storage networking, bare-metal provisioning, or high-speed optical networking.
- Greenfield data centre network build experience or CCNP, CCIE, or equivalent certification.
Culture & Benefits
- Competitive salary with eligibility for equity.
- Biannual bonus scheme.
- Fully expensed technology matched to role requirements.
- Private health insurance.
- Breakfast and dinner allowance for office-based employees.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
Nscale
5 дней назад
Staff Network Engineer (AI)
10 дней назад
Network Engineer (Cisco)
5 дней назад
Senior Network Tooling Engineer
In8inity Telecom
16 часов назад
Lead Network Operations Engineer (ISP)
7 дней назад
Senior Network Engineer - 6 month fixed term contract (AWS)
11 дней назад