Назад
Company hidden
6 дней назад

Network Engineer (AI Infrastructure)

Тип работы
fulltime
Английский
b2
Страна
Armenia
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Network Engineer (AI Infrastructure) (InfiniBand/High-Speed Ethernet): Design and operate ultra-low-latency, high-throughput network fabrics for large-scale GPU clusters and AI training workloads with an accent on InfiniBand, RoCEv2, RDMA, and 200/400GbE switching. Focus on building leaf–spine fabrics, optimizing congestion control and storage networking, automating configuration, and maintaining reliable GPU infrastructure.

Location: Yerevan, Armenia

Company

hirify.global develops infrastructure for large-scale GPU clusters and AI training workloads.

What you will do

  • Design and deploy leaf–spine network topologies for GPU clusters using InfiniBand and high-speed Ethernet.
  • Configure and operate Mellanox/NVIDIA switches, RoCEv2 and InfiniBand fabrics, RDMA, PFC, ECN, DCBX, and congestion-control systems.
  • Implement segmentation and multi-tenancy with VLANs, VRFs, VPCs, SDN overlays, BlueField-3 DPUs, and SR-IOV.
  • Monitor fabric health, troubleshoot AI training pipelines and RDMA collectives, and maintain BGP, EVPN, OSPF, MLAG/VPC, firmware, and routing configurations.
  • Deploy and optimize NVMe-over-Fabrics and integrate high-performance storage platforms including WekaFS, VAST, DDN, Pure, and Ceph NVMe tiers.
  • Build automation, dashboards, alerting, fabric validation, auto-provisioning, and drift-detection tooling with Ansible, Terraform, GitOps, Grafana, and Prometheus.

Requirements

  • Deep expertise in InfiniBand HDR/NDR or 200G/400G Ethernet and strong knowledge of RDMA, RoCEv2, QP states, and congestion management.
  • Advanced routing and switching experience with BGP, EVPN, OSPF, MLAG/VPC, plus strong Linux networking fundamentals.
  • Experience with Mellanox/NVIDIA platforms, including ONIE, Cumulus, MLNX-OS, UFM, or NEO.
  • Experience supporting AI training clusters with H100, B200, A100, or similar GPUs.
  • Familiarity with BlueField DPUs, SR-IOV, VirtIO, OVN, Kube-OVN, or Calico networking.
  • Understanding of NVMe-oF over TCP or RDMA and experience integrating high-performance storage networks.

Nice to have

  • Kubernetes networking experience with CNI, Cilium, Calico, or Multus.
  • OpenStack Neutron or OVN experience.
  • Understanding of GPU virtualization and MIG networking patterns.

Culture & Benefits

  • Full-time role based in Yerevan, Armenia.
  • Collaboration with electrical, mechanical, and compute teams on rack layouts, airflow, cabling, and data-center integration.
  • Work on high-performance AI training and inference infrastructure.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →