Назад
Company hidden
3 дня назад

HPC Specialist (AI/ML)

Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US/Europe/Canada
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
HPC Specialist (AI/ML): Building and operating GPU infrastructure for large-scale LLM inference and ML workloads with an accent on distributed model serving, Kubernetes orchestration, networking, and storage optimization. Focus on troubleshooting performance bottlenecks across hardware and software layers, implementing inference acceleration, and improving reliability through monitoring, capacity planning, and incident response.

Location: Montreal, Canada

Company

hirify.global is a diversified trading firm that combines sophisticated technology and trading expertise across global financial markets, with additional strategies in real estate, venture capital, and cryptoassets.

What you will do

  • Deploy, maintain, and optimize GPU server fleets for large-scale LLM inference workloads.
  • Architect distributed serving solutions for multi-node, multi-GPU model deployments.
  • Manage GPU-enabled Kubernetes clusters and configure networking, load balancers, firewalls, and inter-node communication.
  • Implement storage solutions for model weights and inference caches.
  • Troubleshoot performance bottlenecks across hardware, drivers, networking, and application layers.
  • Collaborate with ML engineers on model profiling, inference acceleration, monitoring, capacity planning, and incident response.

Requirements

  • Bachelor’s or Master’s degree in Computer Science, Systems Engineering, or a related field.
  • 5+ years of experience in DevOps, SRE, or infrastructure engineering.
  • Strong experience with GPU infrastructure, GPU driver management, and model serving frameworks such as vLLM and SGLang.
  • Hands-on experience optimizing deep learning inference or training workloads on GPU clusters.
  • Deep Linux systems knowledge, including networking, storage optimization, and Kubernetes orchestration.
  • Experience with Ansible, Terraform, or similar infrastructure-as-code tools, plus Python and Bash scripting, distributed systems, TCP/IP, HTTP/2, load balancing, Prometheus, and Grafana.

Culture & Benefits

  • Work in an environment emphasizing autonomy, innovation, integrity, curiosity, and challenging consensus.
  • Collaborate with AI, ML, and systematic strategies specialists on infrastructure supporting global trading activities.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →