Назад
Company hidden
8 дней назад

Cluster Administration Engineer (AI)

200 000 - 400 000SGD
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Singapore
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Cluster Administration Engineer (AI): Operating high-performance GPU and HPC clusters that support vLLM development in Singapore with an accent on cluster health, GPU availability, monitoring, scheduling, and diagnostics. Focus on automating infrastructure operations, resolving urgent compute incidents, and scaling cluster provisioning across compute providers.

Location: On-site in Singapore

Annual salary: S$200,000–S$400,000 plus equity

Company

hirify.global develops and operates infrastructure for vLLM, an AI inference engine, with a focus on making model inference faster and more cost-effective.

What you will do

  • Own the health, availability, observability, and usability of high-performance GPU and HPC clusters.
  • Manage GPU availability, monitoring, alerting, scheduling, access, diagnostics, and incident response.
  • Operate GPU servers, troubleshoot node failures, memory errors, driver issues, scheduler problems, and hardware faults.
  • Standardize provisioning, operations, debugging, and scaling across neo-cloud and dedicated compute providers.
  • Automate operational workflows and improve cluster utilization while reducing idle or unavailable GPU capacity.
  • Work with engineering leadership and infrastructure owners to keep compute systems productive for engineering teams.

Requirements

  • Bachelor’s degree or equivalent experience in computer science, engineering, systems administration, or a related field.
  • Hands-on experience administering large compute, HPC, research, supercomputing, or production GPU clusters.
  • Strong Linux systems administration skills, including networking, processes, storage, package management, shell scripting, logs, access control, and debugging.
  • Experience with cluster scheduling and resource allocation using SLURM, Kubernetes, or equivalent tools.
  • Ability to own urgent infrastructure incidents end to end when compute issues block engineering teams.
  • Experience with Bash, Python, Ansible, Terraform, Helm, or similar automation tools.

Nice to have

  • Experience with GPU compute providers such as Lambda, CoreWeave, Crusoe, Nebius, Together, Fireworks, or RunPod.
  • Knowledge of InfiniBand, RoCE, NVLink, NVSwitch, RDMA, NCCL, or equivalent high-performance GPU networking systems.
  • Experience with NFS, Lustre, Ceph, distributed filesystems, or other high-throughput storage systems.
  • Background in research computing, scientific computing, ML infrastructure, SRE, platform engineering, or infrastructure operations.
  • Experience operating Kubernetes for ML or GPU workloads and standardizing infrastructure across multiple providers.

Culture & Benefits

  • Visa sponsorship is available on a case-by-case basis.
  • Medical, dental, and vision coverage.
  • Equity in addition to the annual salary.
  • Infrastructure supports engineering and research workloads around the clock.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →