Назад
Company hidden
23 часа назад

GPU Systems Engineer

200 000 - 300 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
GPU Systems Engineer (GPU infrastructure/HPC): Design, deploy, and operate distributed GPU clusters spanning thousands of nodes, with an accent on Linux systems, GPU workload performance, and infrastructure automation. Focus on profiling AI workloads, troubleshooting GPUDirect RDMA across hardware and network layers, and building self-healing tooling for large-scale compute fleets.

Location: New York, hybrid

Salary: $200,000–$300,000 annual base salary, plus eligible discretionary bonus.

Company

hirify.global is a quantitative trading firm that develops high-performance electronic trading infrastructure and operates a global engineering organization.

What you will do

  • Design, deploy, scale, and operate distributed GPU clusters, including hardware selection and network topology.
  • Identify performance bottlenecks across compute, storage, networking, and their integration points.
  • Profile and benchmark GPU workloads with researchers and convert findings into measurable performance improvements.
  • Build provisioning, monitoring, diagnostics, and self-healing automation for fleets of thousands of nodes.
  • Own infrastructure projects from architecture and implementation through long-term support.
  • Qualify new hardware and software generations and work with vendors to resolve complex issues.

Requirements

  • 5+ years of experience engineering large-scale Linux systems in HPC, AI, or distributed-infrastructure environments.
  • Deep Linux knowledge, including installation, performance tuning, debugging, and kernel-level investigation.
  • Hands-on experience troubleshooting distributed GPU workloads and understanding GPU performance.
  • Working experience with GPUDirect RDMA and data movement between GPUs and networks.
  • Python for automation and tooling, plus CUDA or C/C++ experience for reading, profiling, and debugging GPU code.
  • Experience with configuration management tools such as Salt, Ansible, Puppet, or Chef, and the ability to diagnose issues across hardware, operating-system, and network layers.

Nice to have

  • Experience with additional parts of the NVIDIA stack, including NCCL and NVLink.

Culture & Benefits

  • Hybrid working opportunities in a collaborative, low-hierarchy environment.
  • Generous paid time off policies.
  • Savings plans and financial wellness tools available in each region.
  • Free breakfast, lunch, and snacks in the office.
  • Wellness reimbursements, sports teams, fitness events, volunteer opportunities, and charitable giving.
  • Workshops and continuous learning opportunities, along with regular social events.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →