Назад
7 дней назад

Principal Network Engineer (AI Infrastructure)

Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Principal Network Engineer (AI Infrastructure): Owning and evolving high-speed InfiniBand, RoCE, and RDMA network fabrics for tightly coupled GPU clusters with an accent on reliability, scalability, and performance predictability. Focus on designing large-scale interconnect architectures, resolving complex cross-layer incidents, and defining operational standards for AI infrastructure networking.

Location: Houston, New York, San Francisco, or Seattle, United States

Company

GPU cloud infrastructure engineered for AI startups and enterprise customers, providing high-performance and cost-efficient computing platforms.

What you will do

  • Own the technical direction and operational strategy for AI interconnect networks.
  • Design, review, and evolve large-scale InfiniBand and RoCE fabric architectures.
  • Lead investigations into complex network incidents and implement systemic fixes.
  • Define standards for hardware configuration, congestion control, routing, firmware lifecycle management, and safe changes.
  • Partner with SRE, compute platform, and network architecture teams on end-to-end system design.
  • Mentor network engineers and improve uptime, latency consistency, capacity efficiency, and incident rates.

Requirements

  • 10+ years of network engineering experience focused on HPC, AI, or hyperscale data center networking.
  • Expert operational and architectural experience with InfiniBand and/or large-scale RoCE fabrics.
  • Deep knowledge of RDMA internals, congestion management, and fabric-level failure modes.
  • Strong expertise in data center routing and control planes, including BGP, OSPF, and ECMP.
  • Ability to debug cross-layer issues involving hardware, firmware, kernels, and application communication libraries.
  • Experience leading complex cross-team technical initiatives without direct authority.

Nice to have

  • Production experience with NVIDIA/Mellanox networking platforms in AI or HPC environments.
  • Familiarity with distributed training frameworks and GPU communication patterns.
  • Experience designing observability systems for high-cardinality, high-throughput environments.
  • Experience influencing platform or infrastructure strategy at scale.

Culture & Benefits

  • Collaborative, supportive, and innovation-focused environment.
  • Competitive base salary and equity package with reviews every 12 months.
  • Career progression plan with opportunities to lead, challenge existing approaches, and own impact.
  • Inclusive workplace with support for individual accommodation needs.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →