обновлено 5 дней назад
Software Engineer (GPU Networking)
165 000 - 330 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Software Engineer (GPU Networking): Building and optimizing high-performance GPU networking and distributed systems for AI inference with an accent on integrating RDMA capabilities and co-optimizing communication alongside computation. Focus on architecting the software fabric that unifies thousands of GPUs, enabling serverless-grade startup speeds for LLMs, and deep-diving into bleeding-edge hardware performance.
Location: Hybrid in San Francisco, Montreal, New York, Seattle, or Toronto
Salary: $165K–$330K annually, plus equity
Company
An AI infrastructure company building inference platforms, developer tooling, and distributed hardware systems for mission-critical AI applications.
What you will do
- Integrate RDMA, RoCE, and InfiniBand capabilities into the inference stack.
- Build and tune networking layers for disaggregated KV cache offload and Wide Expert Parallelism.
- Develop checkpointing and storage mechanisms that enable sub-10-second startup for trillion-parameter models.
- Characterize networking performance on H100, H200, B200, B300, GB200, and GB300 NVL72 clusters and create hardware acceptance tests.
- Design observability tools for packet flow, congestion, and effective bandwidth across GPU interconnects.
- Optimize NCCL and NVSHMEM communication and develop custom kernels to overlap computation with data transfer.
Requirements
- Deep experience with high-performance networking protocols, including InfiniBand and RoCE v2.
- Fluency in C++ or Python and the ability to work across high-level software and hardware.
- Strong understanding of memory hierarchies in modern NVIDIA architectures, including H100 and Blackwell GPUs.
- Experience debugging low-level systems such as TensorRT-LLM, custom C++ or Python bindings, or NVLink topologies.
- Ability to evaluate when to use existing solutions and when to build custom networking infrastructure.
Nice to have
- Knowledge of NCCL, NVSHMEM, and UCX.
- Rust experience in systems-level or performance-critical networking.
- Experience with GPUDirect Storage or high-performance filesystems such as Weka or 3FS.
- Familiarity with TensorRT-LLM, vLLM, or SGLang.
- Experience benchmarking and qualifying new hardware clusters.
Culture & Benefits
- Work with bleeding-edge Blackwell and upcoming Rubin GPU architectures.
- Operate across hardware interconnects, communication kernels, and distributed inference strategies.
- Competitive compensation with meaningful equity.
- U.S. employees receive full medical, dental, and vision coverage, plus access to a company-facilitated 401(k).
- Flexible PTO, company-wide winter break, paid parental leave, and a fertility and family-building stipend.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
10 дней назад
Founding GPU Engineer (AI)
11 дней назад
Principal Software Engineer (AI)
160 200 - 425 000$
6 дней назад
Senior AI Engineer Developer
137 000 - 315 000$
10 дней назад
Parallel Computing Engineer (CUDA)
130 000 - 180 000$
8 дней назад
AI Engineer
150 000 - 250 000$
7 дней назад
Software Engineer (AI)
210 000 - 265 000$