Назад
Company hidden
9 дней назад

Senior AI Data Center Network Engineer (InfiniBand)

Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
c1
Страна
Singapore/US/Norway +2 еще
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior AI Data Center Network Engineer (InfiniBand) (AI infrastructure): Architecting, deploying, and operating high-performance networks for large-scale GPU clusters with an accent on InfiniBand, RoCEv2, VXLAN EVPN, and SDN architectures. Focus on troubleshooting RDMA packet loss and congestion, automating infrastructure with Python, Ansible, and Terraform, and maintaining reliable AI training and inference connectivity.

Location: Needham, Massachusetts, United States

Company

hirify.global develops Bitcoin mining infrastructure, AI computational infrastructure, data centers, and cloud capabilities for artificial intelligence workloads.

What you will do

  • Architect scalable, highly available data center networks using Spine-Leaf topology, DCN, DCI, VXLAN EVPN, and SDN.
  • Design, deploy, and optimize GPU cluster networks based on InfiniBand and RoCEv2, including NVIDIA Spectrum and Quantum switches.
  • Investigate complex AI networking issues involving RDMA packet loss, PFC, ECN, congestion, latency, and NCCL communication timeouts.
  • Develop network automation and Infrastructure as Code using Python, Ansible, and Terraform.
  • Monitor fabric health with NVIDIA UFM, NetQ, Zabbix, and Prometheus.
  • Support Kubernetes, Docker, hybrid cloud networking, incident response, capacity expansions, firmware upgrades, and root-cause analysis.

Requirements

  • Bachelor’s degree or higher in Computer Science, Network Engineering, Telecommunications, or a related field.
  • 8–10 years of experience in large-scale data center network operations, architecture, and engineering.
  • Expertise in TCP/IP, BGP, OSPF, ISIS, VXLAN EVPN, and Spine-Leaf data center technologies.
  • Hands-on experience with InfiniBand, Subnet Manager, RoCEv2, PFC, ECN, and congestion control.
  • Experience with Cisco Nexus, NVIDIA/Mellanox, or Juniper switches and routers, plus Linux administration.
  • Professional fluency in English and Chinese is required.

Culture & Benefits

  • Work with the AI Cloud team on infrastructure supporting AI training and inference workloads.
  • Collaborate with global cross-functional teams.
  • Operate large-scale AI infrastructure for customers across multiple markets.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →