6 дней назад
Network Engineer (AI Infrastructure)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Network Engineer (AI Infrastructure) (InfiniBand/High-Speed Ethernet): Design and operate ultra-low-latency, high-throughput network fabrics for large-scale GPU clusters and AI training workloads with an accent on InfiniBand, RoCEv2, RDMA, and 200/400GbE switching. Focus on building leaf–spine fabrics, optimizing congestion control and storage networking, automating configuration, and maintaining reliable GPU infrastructure.
Location: Yerevan, Armenia
Company
develops infrastructure for large-scale GPU clusters and AI training workloads.
What you will do
- Design and deploy leaf–spine network topologies for GPU clusters using InfiniBand and high-speed Ethernet.
- Configure and operate Mellanox/NVIDIA switches, RoCEv2 and InfiniBand fabrics, RDMA, PFC, ECN, DCBX, and congestion-control systems.
- Implement segmentation and multi-tenancy with VLANs, VRFs, VPCs, SDN overlays, BlueField-3 DPUs, and SR-IOV.
- Monitor fabric health, troubleshoot AI training pipelines and RDMA collectives, and maintain BGP, EVPN, OSPF, MLAG/VPC, firmware, and routing configurations.
- Deploy and optimize NVMe-over-Fabrics and integrate high-performance storage platforms including WekaFS, VAST, DDN, Pure, and Ceph NVMe tiers.
- Build automation, dashboards, alerting, fabric validation, auto-provisioning, and drift-detection tooling with Ansible, Terraform, GitOps, Grafana, and Prometheus.
Requirements
- Deep expertise in InfiniBand HDR/NDR or 200G/400G Ethernet and strong knowledge of RDMA, RoCEv2, QP states, and congestion management.
- Advanced routing and switching experience with BGP, EVPN, OSPF, MLAG/VPC, plus strong Linux networking fundamentals.
- Experience with Mellanox/NVIDIA platforms, including ONIE, Cumulus, MLNX-OS, UFM, or NEO.
- Experience supporting AI training clusters with H100, B200, A100, or similar GPUs.
- Familiarity with BlueField DPUs, SR-IOV, VirtIO, OVN, Kube-OVN, or Calico networking.
- Understanding of NVMe-oF over TCP or RDMA and experience integrating high-performance storage networks.
Nice to have
- Kubernetes networking experience with CNI, Cilium, Calico, or Multus.
- OpenStack Neutron or OVN experience.
- Understanding of GPU virtualization and MIG networking patterns.
Culture & Benefits
- Full-time role based in Yerevan, Armenia.
- Collaboration with electrical, mechanical, and compute teams on rack layouts, airflow, cabling, and data-center integration.
- Work on high-performance AI training and inference infrastructure.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
2 часа назад
Network Engineer (AI)
202 000 - 261 000$
Nscale
7 дней назад
Principal Network Engineer (AI Infrastructure)
3 дня назад
Senior Network Engineer (AI Infrastructure)
150 000 - 190 000$
6 дней назад
Sr. Network Engineer (AI)
5 дней назад
WAN Network Engineer
3 дня назад
Senior Network Engineer (InfiniBand / UFM)
170 000 - 210 000$