Назад
Company hidden
17 часов назад

HPC Engineer (AI)

Формат работы
remote (только Europe)
Тип работы
fulltime
Грейд
middle/senior
Английский
b2
Страна
UK/US/Europe +1 еще
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
HPC Engineer (AI): Operating and scaling bare-metal and virtualized GPU clusters for an AI cloud with an accent on InfiniBand fabrics, shared filesystems, and workload orchestration. Focus on tuning cluster performance, troubleshooting the full hardware and networking stack, automating operations, and maintaining reliability for customer workloads.

Location: Fully remote within the EU

Company

hirify.global is building a full-stack AI cloud spanning data centers, hardware, and a cloud platform for AI teams.

What you will do

  • Administer bare-metal and virtualized GPU/HPC clusters from provisioning through ongoing operations.
  • Design, deploy, and tune InfiniBand fabrics, including topology planning, subnet management, and performance validation.
  • Deploy and operate shared and parallel filesystems for training and inference workloads.
  • Troubleshoot hardware and connectivity issues across fibers, transceivers, NICs, switches, drivers, and firmware, working with remote-hands and data center teams.
  • Operate Slurm or equivalent workload schedulers and maintain accurate issue tracking, IPAM, and DCIM records.
  • Participate in on-call rotations, incident response, and integration of new clusters with platform, network, and storage teams.

Requirements

  • Solid Linux administration skills and deep knowledge of InfiniBand clustering, fabric design, subnet management, and performance tuning.
  • Experience with shared filesystems such as Lustre, GPFS/Spectrum Scale, or WekaFS.
  • Knowledge of NCCL, CUDA, DOCA, the NVIDIA stack, and Slurm or comparable workload scheduling solutions.
  • Ability to diagnose and resolve hardware issues remotely with remote-hands teams.
  • Experience operating production clusters where uptime and performance affect customer workloads.
  • Scripting and automation skills with Python, Bash, or Ansible.

Nice to have

  • RoCEv2 or Spectrum-X knowledge.
  • GPU health-checking and diagnostics experience, including DCGM or field diagnostics.
  • Experience with bare-metal provisioning, GPU-aware virtualization, containerization, or observability stacks.
  • Understanding of agentic guardrails for administration of complex systems.

Culture & Benefits

  • Full-time, permanent employment with remote work within the EU.
  • Cash and equity compensation.
  • Healthcare, lunch, wellbeing benefits, and other fringe benefits.
  • Opportunity to work with engineers, researchers, and partners across the global AI ecosystem.
  • Participation in a fast-growing, profitable operation with an internationally diverse workforce.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →