Назад
Company hidden
2 дня назад

Senior Software Developer (AI/HPC)

Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Israel
Вакансия из списка Hirify.GlobalВакансия из Hirify RU Global, списка компаний с восточно-европейскими корнями
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/

TL;DR

Senior Software Developer (AI/HPC): Designing and implementing scale-out communication algorithms for AI hardware accelerators with an accent on topology-aware transport layers and collective performance. Focus on optimizing end-to-end bandwidth and latency across multi-node pods, ensuring numerical determinism, and co-designing interconnect features with hardware teams.

Location: Must be based in Israel (Haifa or Petah-Tikva). This role follows a hybrid work model requiring on-site presence at hirify.global sites.

Company

A global leader in semiconductor design and manufacturing, driving innovation in computing and AI hardware.

What you will do

  • Design and implement collective algorithms tuned to interconnect topology and performance profiles.
  • Build a topology-aware transport layer for tray, rack, and pod-level communication.
  • Optimize end-to-end collective performance across multi-node pods by identifying and resolving bottlenecks.
  • Co-design interconnect features with hardware teams and runtime engineers.
  • Ensure correctness and numerical determinism of reductions across large-scale distributed systems.

Requirements

  • 5+ years of experience in AI, Systems, or HPC software development using C++.
  • Strong expertise in concurrency and lock-free design patterns.
  • Hands-on experience with collective libraries such as NCCL or MPI.
  • Deep understanding of communication patterns for tensor, pipeline, and expert parallelism.
  • Proven skills in performance engineering, including profiling and bandwidth/latency tuning.

Nice to have

  • Experience with topology/placement algorithms and congestion control.
  • Exposure to large-scale distributed training or inference stacks like vLLM, SGLang, or TensorRT-LLM.
  • Knowledge of RDMA, InfiniBand, RoCE, or GPUDirect-style transfers.

Culture & Benefits

  • Hybrid work model offering flexibility between on-site and off-site work.
  • Opportunity to work on cutting-edge AI hardware and large-scale interconnect technologies.
  • Collaborative environment working closely with hardware and runtime engineering pillars.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →