обновлено 24 дня назад
Senior Software Developer - Network and Collectives (C++)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Software Developer - Network and Collectives (C++): Building topology-aware collective communication algorithms and transport layers for an AI accelerator across tray, rack, pod, and cluster interconnects with an accent on multi-node performance, concurrency, and numerical determinism. Focus on profiling bandwidth and latency, co-designing interconnect features with hardware, and optimizing all-reduce, all-gather, reduce-scatter, and point-to-point communication at fabric scale.
Location: Hybrid work model in Israel, with primary location in Haifa and an additional location in Petah-Tikva. The role combines on-site work at the assigned site with off-site work.
Company
develops technology for online grocery commerce and AI accelerator infrastructure.
What you will do
- Design and implement collective communication algorithms tuned to interconnect topology and bandwidth and latency characteristics.
- Build topology-aware transport across tray, rack, and pod interconnects.
- Optimize collective performance across multi-node pods through profiling, bottleneck analysis, and fabric-level tuning.
- Co-design interconnect features with hardware teams and coordinate partitioning, overlap, and scheduling with the Multi-Node Runtime team.
- Own correctness and numerical determinism of reductions at scale.
Requirements
- 5+ years of experience in AI, systems, or HPC software development.
- Strong C++ skills with experience in concurrency and lock-free design.
- Hands-on experience with collective libraries such as NCCL or MPI, or with scale-out communication.
- Working knowledge of tensor, pipeline, and expert parallelism and their mapping to collectives and hardware fabrics.
- Experience with profiling, bandwidth and latency tuning, and roofline reasoning.
Nice to have
- Experience with topology or placement algorithms, congestion control, or multi-tenant fabrics.
- Exposure to large-scale distributed training or inference.
- Experience with vLLM, SGLang, TensorRT-LLM, or DeepSeek.
- Experience with RDMA, InfiniBand, RoCE, GPUDirect-style transfers, or comparable fabric technologies.
Culture & Benefits
- Hybrid work model combining on-site and off-site work.
- Close collaboration with hardware and multi-node runtime teams.
- Shift 1 aligned with Israel.
- Full-time experienced-hire position.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →