Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Principal Network Engineer (AI Infrastructure): Owning and evolving high-speed InfiniBand, RoCE, and RDMA network fabrics for tightly coupled GPU clusters with an accent on reliability, scalability, and performance predictability. Focus on designing large-scale interconnect architectures, resolving complex cross-layer incidents, and defining operational standards for AI infrastructure networking.
Location: Houston, New York, San Francisco, or Seattle, United States
Company
GPU cloud infrastructure engineered for AI startups and enterprise customers, providing high-performance and cost-efficient computing platforms.
What you will do
- Own the technical direction and operational strategy for AI interconnect networks.
- Design, review, and evolve large-scale InfiniBand and RoCE fabric architectures.
- Lead investigations into complex network incidents and implement systemic fixes.
- Define standards for hardware configuration, congestion control, routing, firmware lifecycle management, and safe changes.
- Partner with SRE, compute platform, and network architecture teams on end-to-end system design.
- Mentor network engineers and improve uptime, latency consistency, capacity efficiency, and incident rates.
Requirements
- 10+ years of network engineering experience focused on HPC, AI, or hyperscale data center networking.
- Expert operational and architectural experience with InfiniBand and/or large-scale RoCE fabrics.
- Deep knowledge of RDMA internals, congestion management, and fabric-level failure modes.
- Strong expertise in data center routing and control planes, including BGP, OSPF, and ECMP.
- Ability to debug cross-layer issues involving hardware, firmware, kernels, and application communication libraries.
- Experience leading complex cross-team technical initiatives without direct authority.
Nice to have
- Production experience with NVIDIA/Mellanox networking platforms in AI or HPC environments.
- Familiarity with distributed training frameworks and GPU communication patterns.
- Experience designing observability systems for high-cardinality, high-throughput environments.
- Experience influencing platform or infrastructure strategy at scale.
Culture & Benefits
- Collaborative, supportive, and innovation-focused environment.
- Competitive base salary and equity package with reviews every 12 months.
- Career progression plan with opportunities to lead, challenge existing approaches, and own impact.
- Inclusive workplace with support for individual accommodation needs.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →