3 дня назад
Cluster Administration Engineer (GPU/HPC)
200 000 - 400 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Cluster Administration Engineer (GPU/HPC): Operating high-performance GPU and HPC clusters that support vLLM development with an accent on cluster health, GPU availability, scheduling, monitoring, and incident response. Focus on automating provisioning and diagnostics, improving resource utilization, and debugging complex networking, storage, and compute failures across providers.
Location: San Francisco, California; remote work may be considered within the US for exceptional candidates
Salary: $200,000–$400,000 USD annual salary plus equity
Company
develops and operates infrastructure for vLLM, with a mission to make AI inference cheaper and faster.
What you will do
- Own the health, availability, observability, and usability of high-performance GPU and HPC clusters.
- Monitor GPU availability, cluster scheduling, access, diagnostics, alerting, and incident response.
- Operate GPU servers, troubleshoot node failures, memory errors, driver issues, scheduler problems, and hardware faults.
- Standardize provisioning, operations, debugging, and scaling across neo-cloud and dedicated compute providers.
- Automate operational workflows and improve cluster utilization while reducing idle or unavailable GPU capacity.
- Work with engineering leadership and infrastructure owners to keep compute available for development, testing, and vLLM-related systems.
Requirements
- Bachelor's degree or equivalent experience in computer science, engineering, systems administration, or a related field.
- Hands-on experience administering large compute clusters, HPC environments, research clusters, supercomputing systems, or production GPU clusters.
- Strong Linux administration skills covering networking, processes, storage, package management, shell scripting, logs, access control, and debugging.
- Experience with SLURM, Kubernetes, or equivalent cluster scheduling and resource-allocation tools.
- Ability to own urgent infrastructure incidents end to end and automate workflows with Bash, Python, Ansible, Terraform, Helm, or similar tools.
- On-site work is based in San Francisco; remote candidates must be located in the US.
Nice to have
- Experience with GPU compute providers such as Lambda, CoreWeave, Crusoe, Nebius, Together, Fireworks, or RunPod.
- Knowledge of InfiniBand, RoCE, NVLink, NVSwitch, RDMA, NCCL, NFS, Lustre, Ceph, or other HPC networking and storage systems.
- Experience with secure access, identity, permissions, SSH, VPNs, bastion hosts, secrets, and infrastructure security.
- Background in research computing, scientific computing, ML infrastructure, SRE, platform engineering, or infrastructure operations.
- Experience building monitoring, alerting, runbooks, health checks, remediation workflows, or multi-provider operating standards.
Culture & Benefits
- Hands-on ownership of infrastructure used by engineers and researchers.
- Health, dental, and vision benefits.
- 401(k) company match.
- Equity included in compensation.
- Visa sponsorship is available on a case-by-case basis.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
5 дней назад
Senior HPC Cluster Engineer
145 920 - 209 241$
4 дня назад
Staff Slurm Cluster & HPC Engineer
4 дня назад
Cloud Infrastructure Engineer
150 000 - 167 166$
5 дней назад
Member of Technical Staff - System Engineering
200 000 - 300 000$
5 дней назад
Systems Engineer - R&D (Linux/HPC)
150 000 - 300 000$
5 дней назад