Назад

Не получаете ответ?

Telegram-вакансии старше 7 дней могут быть уже неактуальны.

10 дней назад

HPC Infrastructure Engineer - GPU Clusters

Формат работы
remote (только USA)
Тип работы
fulltime
Страна
US
vacancy_detail.hirify_telegram_tooltipВакансия из Telegram канала -

Мэтч & Сопровод

Покажет вашу совместимость и напишет письмо

Описание вакансии

TL;DR
HPC Infrastructure Engineer - GPU Clusters (GPU infrastructure/Linux/NVIDIA): Operating and improving GPU infrastructure for provisioning, scheduling, monitoring, upgrades, and capacity planning with an accent on node automation, CUDA/NCCL software stacks, high-speed networking, and storage. Focus on tuning Slurm workloads, diagnosing cluster performance issues, evaluating GPU providers, and maintaining hardware and security operations.

HPC Infrastructure Engineer - GPU Clusters

Company

ElevenLabs

Conditions

4 days ago

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will operate and improve GPU infrastructure across provisioning scheduling monitoring upgrades and capacity planning. You will automate node health checks remediation and capacity burn-in, manage the Linux and NVIDIA software stack, tune job scheduling and storage, investigate performance issues, evaluate GPU providers, and handle hardware and security operations.

Requirements

  • Production experience running large-scale Linux server or GPU environments
  • Knowledge of NVIDIA drivers CUDA NCCL and DCGM or deep systems experience
  • Comfort with bare-metal environments server hardware and high-speed networking
  • Python or Bash automation skills
  • Experience with Ansible or Terraform
  • Ability to analyze metrics logs and PromQL data
  • Experience supporting ML training workloads is beneficial
  • Experience evaluating GPU cloud providers is beneficial
  • Experience with parallel filesystems or large-scale object storage is beneficial
  • Experience with BMC IPMI Redfish automation or PXE provisioning is beneficial
  • Awareness of power and cooling for dense GPU deployments is beneficial

Responsibilities

  • Operate and improve the GPU fleet
  • Automate node health checks draining remediation and burn-in pipelines
  • Manage OS images NVIDIA drivers CUDA container runtimes NCCL and high-speed networking
  • Run and tune Slurm or similar job scheduling
  • Build and maintain storage for datasets and checkpoints
  • Diagnose and resolve GPU cluster performance problems
  • Evaluate rented GPU capacity and provider SLAs
  • Rack cable and diagnose hardware when needed
  • Maintain access control network isolation and secrets management

Benefits

  • Annual discretionary professional development stipend
  • Annual discretionary social travel stipend
  • Annual company offsite
  • Monthly coworking stipend

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →

Текст вакансии взят без изменений

Источник -