8 дней назад
Cluster Administration Engineer (AI)
200 000 - 400 000SGD
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Cluster Administration Engineer (AI): Operating high-performance GPU and HPC clusters that support vLLM development in Singapore with an accent on cluster health, GPU availability, monitoring, scheduling, and diagnostics. Focus on automating infrastructure operations, resolving urgent compute incidents, and scaling cluster provisioning across compute providers.
Location: On-site in Singapore
Annual salary: S$200,000–S$400,000 plus equity
Company
develops and operates infrastructure for vLLM, an AI inference engine, with a focus on making model inference faster and more cost-effective.
What you will do
- Own the health, availability, observability, and usability of high-performance GPU and HPC clusters.
- Manage GPU availability, monitoring, alerting, scheduling, access, diagnostics, and incident response.
- Operate GPU servers, troubleshoot node failures, memory errors, driver issues, scheduler problems, and hardware faults.
- Standardize provisioning, operations, debugging, and scaling across neo-cloud and dedicated compute providers.
- Automate operational workflows and improve cluster utilization while reducing idle or unavailable GPU capacity.
- Work with engineering leadership and infrastructure owners to keep compute systems productive for engineering teams.
Requirements
- Bachelor’s degree or equivalent experience in computer science, engineering, systems administration, or a related field.
- Hands-on experience administering large compute, HPC, research, supercomputing, or production GPU clusters.
- Strong Linux systems administration skills, including networking, processes, storage, package management, shell scripting, logs, access control, and debugging.
- Experience with cluster scheduling and resource allocation using SLURM, Kubernetes, or equivalent tools.
- Ability to own urgent infrastructure incidents end to end when compute issues block engineering teams.
- Experience with Bash, Python, Ansible, Terraform, Helm, or similar automation tools.
Nice to have
- Experience with GPU compute providers such as Lambda, CoreWeave, Crusoe, Nebius, Together, Fireworks, or RunPod.
- Knowledge of InfiniBand, RoCE, NVLink, NVSwitch, RDMA, NCCL, or equivalent high-performance GPU networking systems.
- Experience with NFS, Lustre, Ceph, distributed filesystems, or other high-throughput storage systems.
- Background in research computing, scientific computing, ML infrastructure, SRE, platform engineering, or infrastructure operations.
- Experience operating Kubernetes for ML or GPU workloads and standardizing infrastructure across multiple providers.
Culture & Benefits
- Visa sponsorship is available on a case-by-case basis.
- Medical, dental, and vision coverage.
- Equity in addition to the annual salary.
- Infrastructure supports engineering and research workloads around the clock.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
11 дней назад
Senior Kubernetes Platform Engineer (AI)
10 дней назад
Forward Deployed Engineer (Fintech)
140 000 - 220 000$
7 дней назад
Staff Software Engineer (Cloud Infrastructure & AI DevOps)
170 000 - 250 000$
Airwallex
10 дней назад
Senior Cloud Infra Engineer (Cloud Networking)
13 дней назад
Head of Cloud Infrastructure (AWS/Kubernetes)
200 000 - 250 000$
OKX
10 дней назад