Назад
Company hidden
обновлено 1 день назад

Datacenter Network Engineer (AI)

150 000 - 300 000$
Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
middle
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Datacenter Network Engineer (AI) (GPU networking and distributed infrastructure): Designing and operating high-performance networks connecting large GPU clusters, with an accent on Ethernet/RoCE, InfiniBand, reliability, and automation. Focus on diagnosing congestion and packet loss, benchmarking distributed training performance, and building monitoring and safe operational workflows.

Location: San Francisco or remote within the United States

Salary: $150,000–$300,000 plus equity incentives

Company

hirify.global is building an open superintelligence stack that provides AI teams with infrastructure for compute, environments, evaluations, secure sandboxes, training, and deployment.

What you will do

  • Design and deploy scalable datacenter network topologies for GPU training, inference, storage, and management traffic.
  • Configure and operate high-performance Ethernet/RoCE and InfiniBand fabrics with standards for routing, redundancy, and capacity.
  • Automate network provisioning, configuration validation, upgrades, and rollback procedures.
  • Diagnose packet loss, congestion, link failures, and collective communication performance across hosts and switches.
  • Benchmark network performance with infrastructure and ML teams and define measurable acceptance criteria.
  • Build monitoring for port health, errors, utilization, congestion, and fabric topology while improving incident response and runbooks.

Requirements

  • 3+ years of production datacenter networking experience.
  • Strong knowledge of Ethernet, TCP/IP, routing, switching, and redundant network design.
  • Hands-on experience with GPU networking using InfiniBand or RoCE.
  • Experience troubleshooting Linux hosts, NICs, switches, and physical links.
  • Ability to automate network operations with Python, Ansible, or comparable tools.
  • Experience with leaf-spine architectures, BGP, ECMP, VLANs, network segmentation, RDMA, congestion control, and lossless Ethernet.

Nice to have

  • Experience operating 400G/800G networks or large multi-rack GPU clusters.
  • NVIDIA Spectrum or Quantum networking experience.
  • NCCL performance analysis and distributed training troubleshooting experience.
  • Experience with EVPN/VXLAN, SONiC, network source-of-truth systems, network simulation, automated validation, or capacity planning.

Culture & Benefits

  • Work on infrastructure for frontier AI research, training, and inference.
  • Collaborate directly with AI startups and enterprises operating large-scale systems.
  • Partner with datacenter operators and hardware vendors on cabling, optics, deployment readiness, and failure resolution.
  • Cash compensation of $150,000–$300,000 plus equity incentives.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →