6 часов назад
Senior SRE (GPU Infrastructure)
168 000 - 252 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior SRE (GPU Infrastructure) (Linux/AWS/OCI): Operating and scaling production GPU clusters for AI training and inference across on-premises and multi-cloud environments with an accent on Linux performance, high-performance networking, and infrastructure automation. Focus on redesigning clusters for scale, debugging GPU and kernel-level failures, building self-healing systems, and strengthening security and compliance.
Location: Hybrid in Redwood City, California, with remote work available in the US
Salary: $168K–$252K base salary, plus equity
Company
builds unified general intelligence that can generate, understand, and operate in the physical world, with a focus on multimodal AI and vision.
What you will do
- Own production GPU clusters for AI training and inference across on-premises infrastructure, AWS, and OCI.
- Maintain high availability and performance across thousands of NVIDIA and AMD GPUs.
- Redesign infrastructure for greater efficiency, reliability, and scale.
- Tune Linux systems at the OS and kernel level and build automation in Python, Go, or Bash.
- Act as the final escalation point for GPU, networking, and system failures, collaborating with vendors such as NVIDIA.
- Strengthen infrastructure security and support SOC 2 Type 1, SOC 2 Type 2, and ISO certifications.
Requirements
- 5+ years of experience as an SRE, production engineer, or infrastructure engineer in large-scale environments.
- Deep hands-on expertise with Linux, containerized systems, and low-level performance debugging.
- Working experience with Terraform, Airflow, and Ray.
- Strong experience with AWS or OCI.
- Practical experience with high-performance networking, including InfiniBand, RDMA, or RoCE.
- Working knowledge of infrastructure security practices and compliance frameworks such as SOC 2 and ISO.
Nice to have
- Expertise with NVIDIA and AMD GPU tooling, including DCGM and ROCm.
- Experience managing large-scale GPU clusters for AI/ML training or inference.
- Familiarity with Kubernetes or orchestration frameworks such as Ray.
- Expertise in data pipelines and infrastructure.
Culture & Benefits
- Hands-on, close-to-the-metal work in a fast-paced and less-structured environment.
- Opportunity to solve complex low-level GPU, networking, Linux, and kernel problems.
- Base salary of $168K–$252K with equity.
- Work across on-premises and multi-cloud infrastructure at significant scale.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
1 день назад
Infrastructure Operations Engineer (AI)
160 000 - 200 000$
4 дня назад
SRE Engineer II (Cloud/DevOps)
141 000 - 162 000$
2 дня назад
Security Platform Engineer (AI)
160 000 - 180 000$
Lambda
1 день назад
Senior Site Reliability Engineer (Kubernetes)
267 000 - 356 000$
3 дня назад
Senior Systems Engineer (Cloud Infrastructure)
150 000 - 250 000$
4 дня назад
Senior Site Reliability Engineer (AI Infrastructure)
215 000 - 275 000$