9 дней назад
Senior AI Data Center Network Engineer (InfiniBand)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior AI Data Center Network Engineer (InfiniBand) (AI infrastructure): Architecting, deploying, and operating high-performance networks for large-scale GPU clusters with an accent on InfiniBand, RoCEv2, VXLAN EVPN, and SDN architectures. Focus on troubleshooting RDMA packet loss and congestion, automating infrastructure with Python, Ansible, and Terraform, and maintaining reliable AI training and inference connectivity.
Location: Needham, Massachusetts, United States
Company
develops Bitcoin mining infrastructure, AI computational infrastructure, data centers, and cloud capabilities for artificial intelligence workloads.
What you will do
- Architect scalable, highly available data center networks using Spine-Leaf topology, DCN, DCI, VXLAN EVPN, and SDN.
- Design, deploy, and optimize GPU cluster networks based on InfiniBand and RoCEv2, including NVIDIA Spectrum and Quantum switches.
- Investigate complex AI networking issues involving RDMA packet loss, PFC, ECN, congestion, latency, and NCCL communication timeouts.
- Develop network automation and Infrastructure as Code using Python, Ansible, and Terraform.
- Monitor fabric health with NVIDIA UFM, NetQ, Zabbix, and Prometheus.
- Support Kubernetes, Docker, hybrid cloud networking, incident response, capacity expansions, firmware upgrades, and root-cause analysis.
Requirements
- Bachelor’s degree or higher in Computer Science, Network Engineering, Telecommunications, or a related field.
- 8–10 years of experience in large-scale data center network operations, architecture, and engineering.
- Expertise in TCP/IP, BGP, OSPF, ISIS, VXLAN EVPN, and Spine-Leaf data center technologies.
- Hands-on experience with InfiniBand, Subnet Manager, RoCEv2, PFC, ECN, and congestion control.
- Experience with Cisco Nexus, NVIDIA/Mellanox, or Juniper switches and routers, plus Linux administration.
- Professional fluency in English and Chinese is required.
Culture & Benefits
- Work with the AI Cloud team on infrastructure supporting AI training and inference workloads.
- Collaborate with global cross-functional teams.
- Operate large-scale AI infrastructure for customers across multiple markets.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
10 дней назад
Senior Network Engineer (AI)
168 000 - 231 000$
Lambda
10 дней назад
Senior Staff Network Engineer (AI Cloud)
324 000 - 432 000$
Nscale
10 дней назад
Senior Principal Frontend Network Engineer (AI Infrastructure)
3 дня назад
Senior Enterprise Infrastructure Engineer (Azure)
13 дней назад
Network Engineer (AI/HPC)
94 000 - 117 000$
Lambda
10 дней назад
Network Architect (AI Infrastructure)
284 000 - 378 000$