Назад
Company hidden
1 день назад

Senior Network Engineer (InfiniBand / UFM)

170 000 - 210 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Network Engineer (InfiniBand / UFM): Design, deploy, automate, and operate high-performance NVIDIA InfiniBand fabrics for large-scale GPU clusters with an accent on UFM administration, distributed AI workload performance, and scalable data center networking. Focus on optimizing congestion control, adaptive routing, QoS, observability, and automation across AI Factory environments supporting thousands of nodes.

Location: New York, New York, United States; San Francisco, California, United States; or Seattle, Washington, United States. Hybrid work model for office-based teams.

Annual base salary: $170,000–$210,000 USD, plus discretionary bonus, equity, and eligible-role benefits.

Company

hirify.global builds an end-to-end platform for developing, training, and deploying AI systems, combining developer-focused software with large-scale AI Factory compute.

What you will do

  • Design, deploy, and maintain large-scale NVIDIA Quantum and Quantum-2 InfiniBand fabrics for AI and ML GPU clusters.
  • Deploy and administer NVIDIA UFM Enterprise for provisioning, monitoring, telemetry, and fabric health.
  • Troubleshoot performance issues affecting NCCL, MPI, GPUDirect RDMA, and distributed AI training workloads.
  • Implement and optimize fat-tree, Dragonfly+, Clos, and spine-leaf architectures, including congestion control, adaptive routing, QoS, and traffic engineering.
  • Automate network provisioning and observability with Python, Ansible, Git, REST APIs, Infrastructure as Code, Prometheus, and Grafana.
  • Support high availability, firmware lifecycle management, incident response, root cause analysis, capacity planning, and networking standards for AI infrastructure.

Requirements

  • 7+ years of data center networking experience, with 3+ years supporting NVIDIA InfiniBand environments; 10+ years is expected for large-scale data center networking.
  • Hands-on NVIDIA UFM Enterprise experience and experience operating Quantum and Quantum-2 InfiniBand switches.
  • Deep knowledge of InfiniBand Architecture, Subnet Manager, adaptive routing, congestion control, PKeys, LIDs, QPs, VLs, and SLs.
  • Strong Ubuntu/Linux administration, Python and Ansible automation, and Layer 2/Layer 3 networking experience.
  • Experience with BGP, EVPN, VXLAN, spine-leaf architectures, packet captures, tcpdump, Wireshark, and ibdiagnet.
  • Strong troubleshooting, documentation, architectural communication, and collaboration skills, with the ability to design systems scaling to thousands of nodes.

Nice to have

  • Experience with Netris and Terraform.
  • Multi-region backbone design experience.
  • Exposure to bare-metal provisioning systems.
  • Experience in high-growth infrastructure startups.

Culture & Benefits

  • Medical, dental, and vision coverage for employees and eligible dependents.
  • RSUs, 401(k) matching in the U.S., and comprehensive paid leave programs.
  • Unlimited PTO, company holidays, floating holidays, and a two-week company-wide winter break.
  • Paid parental and family leave, wellness and work-from-home stipends, and an annual learning allowance.
  • Four weeks of paid sabbatical leave after four years of service, flexible schedules, and complimentary office meals.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →