3 дня назад
HPC Specialist (AI/ML)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
HPC Specialist (AI/ML): Building and operating GPU infrastructure for large-scale LLM inference and ML workloads with an accent on distributed model serving, Kubernetes orchestration, networking, and storage optimization. Focus on troubleshooting performance bottlenecks across hardware and software layers, implementing inference acceleration, and improving reliability through monitoring, capacity planning, and incident response.
Location: Montreal, Canada
Company
is a diversified trading firm that combines sophisticated technology and trading expertise across global financial markets, with additional strategies in real estate, venture capital, and cryptoassets.
What you will do
- Deploy, maintain, and optimize GPU server fleets for large-scale LLM inference workloads.
- Architect distributed serving solutions for multi-node, multi-GPU model deployments.
- Manage GPU-enabled Kubernetes clusters and configure networking, load balancers, firewalls, and inter-node communication.
- Implement storage solutions for model weights and inference caches.
- Troubleshoot performance bottlenecks across hardware, drivers, networking, and application layers.
- Collaborate with ML engineers on model profiling, inference acceleration, monitoring, capacity planning, and incident response.
Requirements
- Bachelor’s or Master’s degree in Computer Science, Systems Engineering, or a related field.
- 5+ years of experience in DevOps, SRE, or infrastructure engineering.
- Strong experience with GPU infrastructure, GPU driver management, and model serving frameworks such as vLLM and SGLang.
- Hands-on experience optimizing deep learning inference or training workloads on GPU clusters.
- Deep Linux systems knowledge, including networking, storage optimization, and Kubernetes orchestration.
- Experience with Ansible, Terraform, or similar infrastructure-as-code tools, plus Python and Bash scripting, distributed systems, TCP/IP, HTTP/2, load balancing, Prometheus, and Grafana.
Culture & Benefits
- Work in an environment emphasizing autonomy, innovation, integrity, curiosity, and challenging consensus.
- Collaborate with AI, ML, and systematic strategies specialists on infrastructure supporting global trading activities.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
4 дня назад
AI Infrastructure Solutions Engineer
4 дня назад
Staff Slurm Cluster & HPC Engineer
5 дней назад
Member of Technical Staff - System Engineering
200 000 - 300 000$
5 дней назад
Colo LL Strategic Specialist - Compute (Ultra Low Latency)
120 000$
5 дней назад
HPC Engineer (AI)
5 дней назад
Senior HPC Cluster Engineer
145 920 - 209 241$