Назад
Company hidden
11 часов назад

Network Platform Engineering Lead (AI)

260 000 - 330 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
lead
Английский
b2
Страна
UK/US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Network Platform Engineering Lead (AI) (RoCE v2/InfiniBand): Building and operating large-scale GPU network fabrics and the software platform that exposes reliable, isolated networking to AI workloads with an accent on Ethernet fabric design, RoCE v2, InfiniBand, and network automation. Focus on designing a single multi-site reference architecture, integrating networking with Kubernetes and IaaS control planes, and leading hands-on reliability, observability, and incident response.

Location: Palo Alto, California, United States; hybrid work with on-site presence required during cluster bring-up when needed

Salary: $260,000–$330,000 per year in California, plus equity and bonus

Company

hirify.global is a vertically integrated AI infrastructure platform building and operating large-scale GPU compute infrastructure, including capital, clusters, networking, and software.

What you will do

  • Lead a network platform engineering team through technical direction, design reviews, code reviews, delivery management, hiring, onboarding, and performance conversations.
  • Define the architecture for compute and storage fabrics, overlays, multi-tenant isolation, edge connectivity, telemetry, and automation across sites.
  • Own fabric standards and reference architecture for Ethernet, RoCE v2, and InfiniBand deployments.
  • Expose network capabilities through APIs, abstractions, and IaaS control-plane integrations.
  • Work with cluster bring-up teams, OEMs, partners, customers, security engineering, and other platform leads to improve designs and operational outcomes.
  • Stay hands-on with complex engineering work, reliability, observability, incident response, post-mortems, and structural remediation.

Requirements

  • 5+ years of large-scale data center or cloud network engineering experience, including at least 2 years leading an engineering team.
  • Production software development experience in Python or Go, with familiarity with Rust and standard review and CI practices.
  • Deep production experience with Ethernet fabrics, leaf-spine architectures, BGP, unnumbered BGP, ECMP, and low-latency operations.
  • Production experience with RoCE v2, PFC, ECN, DCQCN, InfiniBand, fat-tree topologies, UFM, fabric partitioning, adaptive routing, and congestion control.
  • Experience with EVPN, VXLAN, multi-tenant overlays, network automation, configuration as code, source-of-truth systems, CI validation, Kubernetes networking, and GPU infrastructure.
  • Proven people management, multi-vendor network design evaluation, security awareness, stakeholder communication, and willingness to work on site during cluster bring-up.

Nice to have

  • AI-assisted development and agent-assisted engineering workflows.
  • ASN operations, IPv6, NVIDIA Spectrum-X, NetQ, Cumulus, SONiC, whitebox platforms, gNMI, OpenConfig, or NETCONF/YANG.
  • OVN/OVS, SR-IOV, DPDK, BlueField DPU, NVLink, NVSwitch, or NCCL experience.
  • Distributed work across time zones, open-source networking contributions, or related infrastructure projects.

Culture & Benefits

  • Supportive, trusted environment focused on growth, impact, and work-life balance.
  • Equity participation in hirify.global.
  • Retirement or pension contributions.
  • Comprehensive health, wellbeing, and insurance benefits.
  • Generous annual vacation allowance.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →