Назад
Company hidden
15 часов назад

Software Engineer – AI/HPC Middleware

Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
middle
Английский
b2
Страна
US/CR
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Software Engineer – AI/HPC Middleware (C/C++/Linux): Developing and optimizing HPC and AI communication middleware, including MPI, SHMEM, NCCL, and RCCL, over high-performance interconnects with an accent on transport integration, profiling, and distributed systems performance. Focus on implementing upstream open-source contributions, diagnosing customer performance issues, and validating end-to-end behavior across drivers, switches, and application stacks.

Location: Remote for employees residing within the United States; New York, New York is listed in the department location.

Company

hirify.global develops scale-out networking solutions for AI and HPC datacenters, integrating hardware, software, and system-level technologies for GPU, CPU, and accelerator clusters.

What you will do

  • Implement and optimize HPC middleware such as MPI and SHMEM, as well as AI collective communication libraries including NCCL, RCCL, and related CCL stacks.
  • Integrate middleware capabilities with transport and provider layers such as libfabric/OFI, UCX, and verbs-like interfaces.
  • Prepare and submit upstream contributions to MPI, SHMEM, CCL, and related open-source projects.
  • Collaborate with kernel, driver, and switch teams to validate end-to-end performance.
  • Analyze performance traces, reproduce customer issues, and translate findings into fixes.

Requirements

  • 2+ years of systems programming experience in C/C++ on Linux.
  • Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent experience.
  • Exposure to HPC middleware, AI collective communication libraries, CUDA/ROCm, or other CCL stacks through work, coursework, or projects.
  • Familiarity with performance analysis and profiling or tracing tools, or a strong willingness to learn.
  • Foundational understanding of networking and/or RDMA concepts.
  • Residence within the United States is required for this remote position.

Nice to have

  • Experience contributing to open-source projects or working with libfabric providers.
  • Familiarity with Ultra Ethernet specifications, RoCEv2, congestion control, or Ethernet-based RDMA.
  • Experience with benchmarking, profiling, and optimization.
  • Background with Omni-Path/OPX or Ethernet-based HPC fabrics.

Culture & Benefits

  • Flexible remote work environment within the United States.
  • Competitive compensation package including equity, cash, and performance-based incentives.
  • Medical, dental, vision, disability, life, dependent care, accidental injury, and pet insurance.
  • Paid holidays, sick time, bonding leave, pregnancy disability leave, and Open Time Off for eligible regular full-time exempt employees.
  • 401(k) plan with company match.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →