1 день назад
Senior Site Reliability Engineer (Kubernetes)
267 000 - 356 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Site Reliability Engineer (Kubernetes): Operating and scaling bare-metal Kubernetes clusters and building control plane services for AI/ML workloads with an accent on cluster lifecycle automation, platform reliability, and incident response. Focus on designing custom controllers, defining SLOs and SLIs, and solving operational challenges across large-scale GPU and cloud infrastructure.
Location: Hybrid, with presence in the San Francisco, San Jose, or Bellevue office 4 days per week; work from home on Tuesdays
Salary: $267,000–$356,000 annually in San Francisco or San Jose; $240,000–$320,000 annually in Bellevue
Company
Lambda builds AI cloud infrastructure and managed Kubernetes platforms for AI and machine learning workloads.
What you will do
- Operate and maintain bare-metal Kubernetes clusters scaling to thousands of nodes.
- Handle cluster degradation, recovery, resizing, and critical incident response through fleet management tools and an on-call rotation.
- Design and maintain control plane services, Kubernetes operators, and custom controllers.
- Build automation for cluster provisioning, upgrades, patching, and deletion.
- Develop Python and Go tooling to validate platform quality and define SLOs and SLIs for Kubernetes services and workloads.
- Assist customers with Kubernetes workloads, storage, authentication, and integration issues while collaborating with HPC Ops and Datacenter Ops.
Requirements
- 6+ years of experience in SRE, operations engineering, or a similar role.
- Deep experience running Linux systems and production Kubernetes clusters, including on-premises, EKS, GKE, or similar environments.
- Strong programming skills in Go and Python, plus experience with GitOps, Helm, and Kubernetes operators.
- Experience provisioning Kubernetes with kubeadm, Cluster API, or similar tools.
- Familiarity with Prometheus, Grafana, FluentBit, and CI/CD pipelines.
- Ability to work independently, collaborate with engineering teams, and support customers during incidents.
Nice to have
- Expertise with Kubernetes CRDs, CSI, CNI, and operator coding.
- Experience with HPC clusters, AI/ML workloads, large-scale GPU clusters, or hybrid and multi-cloud Kubernetes environments.
- Contributions to CNCF projects or Kubernetes SIGs.
Culture & Benefits
- Opportunity to influence the platform roadmap and reliability practices.
- Collaboration with experienced engineers, plus opportunities to mentor and grow.
- Cash and equity compensation.
- Health, dental, and vision coverage for employees and dependents.
- Wellness and commuter stipends for select roles, a 401(k) plan with a 2% company match for US employees, and flexible paid time off.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
4 дня назад
Senior Site Reliability Engineer (AI Infrastructure)
215 000 - 275 000$
1 день назад
Senior Software Engineer, Platform Infrastructure (AWS/Kubernetes)
190 000 - 220 000$
3 дня назад
Build/Release Engineer (Robotics)
190 000 - 230 000$
6 часов назад
Senior SRE (GPU Infrastructure)
168 000 - 252 000$
Baseten
4 дня назад
Software Engineer (Platform)
165 000 - 330 000$
4 дня назад
Senior Site Reliability Engineer (SRE)
170 000 - 196 000$