обновлено 11 дней назад
Principal Network Engineer (AI Infrastructure)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Principal Network Engineer (AI Infrastructure) (RoCEv2/RDMA): Designing and operating large-scale data center networks for AI and ML clusters ranging from thousands to more than 100,000 GPUs with an accent on congestion management, high-speed Ethernet fabrics, and production reliability. Focus on defining network architecture, tuning PFC, ECN, and DCQCN, and building automation and observability that eliminate manual operations at exascale.
Location: Remote within the United States; employment requires authorization to work in the United States.
Company
TensorWave provides a cloud platform for seamless, secure, reliable, and resilient AI compute at scale.
What you will do
- Own end-to-end network architecture and make high-impact technical decisions across engineering teams.
- Define and standardize large-scale RoCEv2 data center networks for AI and ML clusters from thousands to more than 100,000 GPUs.
- Set and validate congestion management strategies across RDMA fabrics using PFC, ECN, and DCQCN.
- Establish automation, validation, and observability patterns that prevent misconfiguration and reduce manual operations.
- Serve as the technical escalation point for complex failures, scaling limits, and architectural tradeoffs in always-on, multi-tenant environments.
- Remain hands-on with high-speed optics, switching, routing, and congestion management in production clusters.
Requirements
- Bachelor’s degree in Computer Science, Electrical Engineering, or a related technical field, or equivalent practical experience.
- Deep experience designing and operating RDMA and RoCEv2 networks in large-scale production AI or HPC environments.
- Expert knowledge of switching hardware and network operating systems, including Arista, Juniper, and SONiC.
- Hands-on experience with PFC, ECN, DCQCN, high-speed Ethernet fabrics, and performance tuning.
- Experience with 400G and 800G optics, AEC, AOC, DAC, and structured cabling at scale.
- Experience with Python, Ansible, Terraform, Git, and production observability tooling.
Culture & Benefits
- Stock options.
- 100% employer-paid medical, dental, and vision insurance for employees.
- Health Savings Account contributions, Flexible Spending Account, and supplementary health benefits.
- Employer-paid short-term and long-term disability insurance, life insurance options, and Employee Assistance Program.
- Flexible PTO, paid holidays, and parental leave.
- 401(k) and other in-office perks.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
12 дней назад
Senior Network Engineer (AI)
200 000 - 240 000$
13 дней назад
Systems Engineer
14 часов назад
Datacenter Networks Engineer (Palo Alto)
12 дней назад
Senior Network Operations Engineer - DevOps (Network & Cloud)
13 дней назад
Network Team Lead
150 000GBP
12 дней назад