Назад
Company hidden
15 часов назад

Senior Software Engineer (AI/HPC Middleware)

Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US/CR
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Software Engineer (AI/HPC Middleware): Developing and optimizing HPC and AI communication middleware, including MPI, SHMEM, and collective communication libraries, over high-performance interconnects with an accent on low-latency communication, transport integration, and upstream open-source contributions. Focus on profiling and tuning distributed communication paths, validating cluster-scale performance, and solving customer issues across drivers, switches, systems, and application frameworks.

Location: Remote for employees residing within the United States; the position is associated with New York, New York.

Company

hirify.global develops high-performance scale-out networking solutions for AI and HPC datacenters, integrating silicon, software, and system technologies for GPU, CPU, and accelerator clusters.

What you will do

  • Design, develop, and optimize HPC middleware such as MPI and SHMEM, along with AI collective communication libraries including NCCL and RCCL.
  • Integrate middleware capabilities with transport and provider layers such as OFI/libfabric, UCX, and verbs-style interfaces.
  • Prepare and submit upstream contributions to MPI, SHMEM, collective communication, and related open-source projects.
  • Collaborate with kernel, driver, switch, systems, and application teams to deliver end-to-end performance.
  • Analyze performance traces, reproduce customer issues, and translate findings into fixes for real-world workloads.

Requirements

  • 6+ years of systems programming experience in C/C++ on Linux and a bachelor's degree in computer science, engineering, or a related field, or equivalent experience.
  • Hands-on experience with HPC middleware internals such as Open MPI, MPICH, MVAPICH, or SHMEM, or with AI collective communications including NCCL, RCCL, CUDA, or ROCm.
  • Experience delivering low-latency and high-throughput communication paths and diagnosing performance issues with profiling and tracing tools.
  • Working knowledge of OFI/libfabric, UCX, verbs-style concepts, RDMA, and performance tuning.
  • Experience contributing to open-source projects.
  • Applicants must reside within the United States.

Nice to have

  • Experience developing or maintaining libfabric providers.
  • Familiarity with Ultra Ethernet specifications, RoCEv2, congestion control, or Ethernet-based RDMA.
  • Experience with cluster-scale benchmarking, profiling, and optimization.
  • Background with HPC fabrics such as Omni-Path.

Culture & Benefits

  • Flexible remote work environment within the United States.
  • Competitive compensation including base pay, equity, cash, and performance-based incentives.
  • Medical, dental, vision, disability, life, dependent care, accidental injury, and pet insurance.
  • Paid holidays, sick time, bonding leave, pregnancy disability leave, and Open Time Off for eligible full-time exempt employees.
  • 401(k) plan with company match.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →