Назад
Company hidden
3 дня назад

Cluster Administration Engineer (GPU/HPC)

200 000 - 400 000$
Формат работы
remote (только USA)/onsite
Тип работы
fulltime
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Cluster Administration Engineer (GPU/HPC): Operating high-performance GPU and HPC clusters that support vLLM development with an accent on cluster health, GPU availability, scheduling, monitoring, and incident response. Focus on automating provisioning and diagnostics, improving resource utilization, and debugging complex networking, storage, and compute failures across providers.

Location: San Francisco, California; remote work may be considered within the US for exceptional candidates

Salary: $200,000–$400,000 USD annual salary plus equity

Company

hirify.global develops and operates infrastructure for vLLM, with a mission to make AI inference cheaper and faster.

What you will do

  • Own the health, availability, observability, and usability of high-performance GPU and HPC clusters.
  • Monitor GPU availability, cluster scheduling, access, diagnostics, alerting, and incident response.
  • Operate GPU servers, troubleshoot node failures, memory errors, driver issues, scheduler problems, and hardware faults.
  • Standardize provisioning, operations, debugging, and scaling across neo-cloud and dedicated compute providers.
  • Automate operational workflows and improve cluster utilization while reducing idle or unavailable GPU capacity.
  • Work with engineering leadership and infrastructure owners to keep compute available for development, testing, and vLLM-related systems.

Requirements

  • Bachelor's degree or equivalent experience in computer science, engineering, systems administration, or a related field.
  • Hands-on experience administering large compute clusters, HPC environments, research clusters, supercomputing systems, or production GPU clusters.
  • Strong Linux administration skills covering networking, processes, storage, package management, shell scripting, logs, access control, and debugging.
  • Experience with SLURM, Kubernetes, or equivalent cluster scheduling and resource-allocation tools.
  • Ability to own urgent infrastructure incidents end to end and automate workflows with Bash, Python, Ansible, Terraform, Helm, or similar tools.
  • On-site work is based in San Francisco; remote candidates must be located in the US.

Nice to have

  • Experience with GPU compute providers such as Lambda, CoreWeave, Crusoe, Nebius, Together, Fireworks, or RunPod.
  • Knowledge of InfiniBand, RoCE, NVLink, NVSwitch, RDMA, NCCL, NFS, Lustre, Ceph, or other HPC networking and storage systems.
  • Experience with secure access, identity, permissions, SSH, VPNs, bastion hosts, secrets, and infrastructure security.
  • Background in research computing, scientific computing, ML infrastructure, SRE, platform engineering, or infrastructure operations.
  • Experience building monitoring, alerting, runbooks, health checks, remediation workflows, or multi-provider operating standards.

Culture & Benefits

  • Hands-on ownership of infrastructure used by engineers and researchers.
  • Health, dental, and vision benefits.
  • 401(k) company match.
  • Equity included in compensation.
  • Visa sponsorship is available on a case-by-case basis.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →