Назад
обновлено 11 дней назад

Principal Network Engineer (AI Infrastructure)

Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Principal Network Engineer (AI Infrastructure) (RoCEv2/RDMA): Designing and operating large-scale data center networks for AI and ML clusters ranging from thousands to more than 100,000 GPUs with an accent on congestion management, high-speed Ethernet fabrics, and production reliability. Focus on defining network architecture, tuning PFC, ECN, and DCQCN, and building automation and observability that eliminate manual operations at exascale.

Location: Remote within the United States; employment requires authorization to work in the United States.

Company

TensorWave provides a cloud platform for seamless, secure, reliable, and resilient AI compute at scale.

What you will do

  • Own end-to-end network architecture and make high-impact technical decisions across engineering teams.
  • Define and standardize large-scale RoCEv2 data center networks for AI and ML clusters from thousands to more than 100,000 GPUs.
  • Set and validate congestion management strategies across RDMA fabrics using PFC, ECN, and DCQCN.
  • Establish automation, validation, and observability patterns that prevent misconfiguration and reduce manual operations.
  • Serve as the technical escalation point for complex failures, scaling limits, and architectural tradeoffs in always-on, multi-tenant environments.
  • Remain hands-on with high-speed optics, switching, routing, and congestion management in production clusters.

Requirements

  • Bachelor’s degree in Computer Science, Electrical Engineering, or a related technical field, or equivalent practical experience.
  • Deep experience designing and operating RDMA and RoCEv2 networks in large-scale production AI or HPC environments.
  • Expert knowledge of switching hardware and network operating systems, including Arista, Juniper, and SONiC.
  • Hands-on experience with PFC, ECN, DCQCN, high-speed Ethernet fabrics, and performance tuning.
  • Experience with 400G and 800G optics, AEC, AOC, DAC, and structured cabling at scale.
  • Experience with Python, Ansible, Terraform, Git, and production observability tooling.

Culture & Benefits

  • Stock options.
  • 100% employer-paid medical, dental, and vision insurance for employees.
  • Health Savings Account contributions, Flexible Spending Account, and supplementary health benefits.
  • Employer-paid short-term and long-term disability insurance, life insurance options, and Employee Assistance Program.
  • Flexible PTO, paid holidays, and parental leave.
  • 401(k) and other in-office perks.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →