Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Network Architect (AI Infrastructure) (InfiniBand/Ethernet): Architecting high-performance networking for AI cloud platforms, GPU clusters, storage backends, and multi-tenant environments with an accent on ultra-low latency, high bandwidth, and fault tolerance. Focus on evaluating next-generation fabrics, designing BGP/EVPN architectures, optimizing GPU and storage traffic, and enabling network automation and telemetry.
Location: Hybrid, with presence required four days per week at the San Jose, San Francisco, or Bellevue office; Tuesday is the designated work-from-home day.
Salary: $315,000–$420,000 annually for San Francisco/San Jose; $284,000–$378,000 annually for Bellevue.
Company
Lambda builds AI cloud infrastructure for AI researchers, enterprises, and hyperscalers, with a focus on large-scale GPU computing.
What you will do
- Architect high-performance, ultra-low-latency, high-bandwidth networks for cloud platforms.
- Define network topologies and architectural patterns for GPU clusters, storage backends, and multi-tenant environments.
- Evaluate and benchmark InfiniBand, RoCE, high-speed Ethernet, Ultra Ethernet, and other next-generation networking technologies.
- Develop network architecture standards, reference designs, and scalability roadmaps for multi-site and hybrid environments.
- Partner with compute and storage architects to provide resilient end-to-end data flow and fault tolerance.
- Guide network automation, provisioning, telemetry, troubleshooting, and operational visibility while mentoring engineers and cross-functional teams.
Requirements
- 7+ years of experience architecting high-performance data center networks, preferably for HPC, AI/ML, or large-scale cloud infrastructure.
- Deep expertise in InfiniBand HDR/NDR, advanced Ethernet fabrics, RoCE, and RDMA.
- Strong knowledge of switching architectures, congestion control, QoS, VXLAN, EVPN, BGP-based fabrics, eBGP underlays, MP-BGP EVPN, ECMP, route policy, and convergence.
- Experience designing low-latency, high-throughput, redundant, fault-tolerant networks for GPU-to-GPU and storage traffic.
- Expertise in optical networking, including transceivers, fiber, link budgets, breakout architectures, and WDM.
- Experience with SDN, APIs, automation, network telemetry, and cross-functional technical leadership.
Nice to have
- AI workload profiling, NCCL, MPI, and distributed-training network tuning experience.
- Experience with network automation frameworks and telemetry tools.
- Exposure to DPU/SmartNIC technologies such as NVIDIA BlueField.
- Knowledge of multi-site interconnects, DWDM, and metro/long-haul networking.
- Experience designing customized networks with hyperscale or enterprise customers.
Culture & Benefits
- Cash and equity compensation.
- Health, dental, and vision coverage for employees and dependents.
- Wellness and commuter stipends for select roles.
- 401(k) plan with a 2% company match for USA employees.
- Flexible paid time off.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
11 дней назад
Network Engineer (AI/HPC)
94 000 - 117 000$
8 дней назад
Senior Network Engineer (AI)
168 000 - 231 000$
xAI
9 дней назад
Network Engineer
150 000 - 250 000$
TensorWave
14 дней назад
Principal Network Engineer (AI Infrastructure)
14 дней назад
Network Engineer IV (Cisco ACI)
108 253 - 158 733$
14 дней назад