4 дня назад
Staff AI Infrastructure Engineer (AI)
241 000 - 331 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Staff AI Infrastructure Engineer (AI/HPC): Building and operating multi-site GPU clusters running Slurm on Kubernetes to support frontier-scale AI biology research with an accent on reliability, distributed systems, and high-performance computing. Focus on debugging multi-node training failures, scaling storage and networking infrastructure, automating cluster operations, and improving incident resilience across thousands of GPUs.
Location: Redwood City, California, United States; hybrid with onsite attendance at least 60% of the working month, approximately 3 days per week
Salary: $241,000–$331,000 base pay per year
Company
is a nonprofit open-science research lab building AI, biological foundation models, and laboratory capabilities to accelerate scientific discovery and help cure disease.
What you will do
- Own reliability, observability, and incident response for multi-site GPU clusters running Slurm on Kubernetes.
- Debug infrastructure failures across storage, networking, scheduling, and GPU compute layers.
- Design and validate cluster scaling plans for larger multi-node and multi-thousand-GPU training runs.
- Build automation for capacity planning, GPU utilization monitoring, workload-manager policies, pod lifecycle management, and configuration as code.
- Collaborate with AI researchers and hero-run leads to design infrastructure for frontier-scale workloads.
- Coordinate technical vendor escalations, improve operational resilience, and develop scalable runbooks.
Requirements
- 8+ years of AI/ML infrastructure engineering experience with deep expertise in HPC/Slurm operations, Kubernetes at scale, distributed systems debugging, or GPU infrastructure.
- Strong Linux systems knowledge, including TCP/IP, InfiniBand, RDMA, storage systems, kernel internals, cgroups, namespaces, eBPF, and sysctls.
- Hands-on Kubernetes and cloud-native experience, including pod lifecycle, CNI plugins, StatefulSets, Helm, and GitOps tooling.
- Experience with HPC workload managers, especially Slurm, and proficiency in Python and Bash.
- Experience with observability tools such as Prometheus or VictoriaMetrics, Grafana, DCGM metrics, and distributed tracing.
- Ability to work onsite in Redwood City at least 60% of the working month.
Nice to have
- Go, Rust, or C/C++ experience.
- Experience with Cilium, ArgoCD, or advanced Slurm patterns.
- Distributed AI training experience with NCCL, PyTorch DDP, multi-node job debugging, and checkpoint/restart workflows.
Culture & Benefits
- Open-science and open-source AI research focused on improving human health.
- Collaborative, team-oriented environment with regular in-person work.
- Employer match on 401(k) contributions.
- Paid time off for volunteering.
- Funding for select family-forming benefits and relocation support.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
4 дня назад
Technical Staff
200 000 - 350 000$
4 дня назад
AI Engineer - Cloud Infrastructure (AI)
175 000 - 275 000$
2 дня назад
HPC / AI Software Infrastructure Lead (E)
151 100 - 256 900$
Baseten
4 дня назад
Forward Deployed Engineers (AI)
200 000 - 400 000$
3 дня назад
Staff Software Engineer (AI)
150 000 - 250 000$
3 дня назад
Staff Software Engineer (AI)
215 000 - 250 000$