17 часов назад
HPC Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
HPC Engineer (AI): Operating and scaling bare-metal and virtualized GPU clusters for an AI cloud with an accent on InfiniBand fabrics, shared filesystems, and workload orchestration. Focus on tuning cluster performance, troubleshooting the full hardware and networking stack, automating operations, and maintaining reliability for customer workloads.
Location: Fully remote within the EU
Company
is building a full-stack AI cloud spanning data centers, hardware, and a cloud platform for AI teams.
What you will do
- Administer bare-metal and virtualized GPU/HPC clusters from provisioning through ongoing operations.
- Design, deploy, and tune InfiniBand fabrics, including topology planning, subnet management, and performance validation.
- Deploy and operate shared and parallel filesystems for training and inference workloads.
- Troubleshoot hardware and connectivity issues across fibers, transceivers, NICs, switches, drivers, and firmware, working with remote-hands and data center teams.
- Operate Slurm or equivalent workload schedulers and maintain accurate issue tracking, IPAM, and DCIM records.
- Participate in on-call rotations, incident response, and integration of new clusters with platform, network, and storage teams.
Requirements
- Solid Linux administration skills and deep knowledge of InfiniBand clustering, fabric design, subnet management, and performance tuning.
- Experience with shared filesystems such as Lustre, GPFS/Spectrum Scale, or WekaFS.
- Knowledge of NCCL, CUDA, DOCA, the NVIDIA stack, and Slurm or comparable workload scheduling solutions.
- Ability to diagnose and resolve hardware issues remotely with remote-hands teams.
- Experience operating production clusters where uptime and performance affect customer workloads.
- Scripting and automation skills with Python, Bash, or Ansible.
Nice to have
- RoCEv2 or Spectrum-X knowledge.
- GPU health-checking and diagnostics experience, including DCGM or field diagnostics.
- Experience with bare-metal provisioning, GPU-aware virtualization, containerization, or observability stacks.
- Understanding of agentic guardrails for administration of complex systems.
Culture & Benefits
- Full-time, permanent employment with remote work within the EU.
- Cash and equity compensation.
- Healthcare, lunch, wellbeing benefits, and other fringe benefits.
- Opportunity to work with engineers, researchers, and partners across the global AI ecosystem.
- Participation in a fast-growing, profitable operation with an internationally diverse workforce.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
15 часов назад
Senior HPC Cluster Engineer
145 920 - 209 241$
8 часов назад
Staff Slurm Cluster & HPC Engineer
6 дней назад
Senior Network Engineer (InfiniBand / UFM)
170 000 - 210 000$
Lambda
5 дней назад
Site Reliability Engineer (AI Infrastructure)
240 000 - 356 000$
8 часов назад
Senior GPU Systems & Fabric Engineer (AI)
11 часов назад