Назад
1 день назад

Senior Site Reliability Engineer (Kubernetes)

267 000 - 356 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Site Reliability Engineer (Kubernetes): Operating and scaling bare-metal Kubernetes clusters and building control plane services for AI/ML workloads with an accent on cluster lifecycle automation, platform reliability, and incident response. Focus on designing custom controllers, defining SLOs and SLIs, and solving operational challenges across large-scale GPU and cloud infrastructure.

Location: Hybrid, with presence in the San Francisco, San Jose, or Bellevue office 4 days per week; work from home on Tuesdays

Salary: $267,000–$356,000 annually in San Francisco or San Jose; $240,000–$320,000 annually in Bellevue

Company

Lambda builds AI cloud infrastructure and managed Kubernetes platforms for AI and machine learning workloads.

What you will do

  • Operate and maintain bare-metal Kubernetes clusters scaling to thousands of nodes.
  • Handle cluster degradation, recovery, resizing, and critical incident response through fleet management tools and an on-call rotation.
  • Design and maintain control plane services, Kubernetes operators, and custom controllers.
  • Build automation for cluster provisioning, upgrades, patching, and deletion.
  • Develop Python and Go tooling to validate platform quality and define SLOs and SLIs for Kubernetes services and workloads.
  • Assist customers with Kubernetes workloads, storage, authentication, and integration issues while collaborating with HPC Ops and Datacenter Ops.

Requirements

  • 6+ years of experience in SRE, operations engineering, or a similar role.
  • Deep experience running Linux systems and production Kubernetes clusters, including on-premises, EKS, GKE, or similar environments.
  • Strong programming skills in Go and Python, plus experience with GitOps, Helm, and Kubernetes operators.
  • Experience provisioning Kubernetes with kubeadm, Cluster API, or similar tools.
  • Familiarity with Prometheus, Grafana, FluentBit, and CI/CD pipelines.
  • Ability to work independently, collaborate with engineering teams, and support customers during incidents.

Nice to have

  • Expertise with Kubernetes CRDs, CSI, CNI, and operator coding.
  • Experience with HPC clusters, AI/ML workloads, large-scale GPU clusters, or hybrid and multi-cloud Kubernetes environments.
  • Contributions to CNCF projects or Kubernetes SIGs.

Culture & Benefits

  • Opportunity to influence the platform roadmap and reliability practices.
  • Collaboration with experienced engineers, plus opportunities to mentor and grow.
  • Cash and equity compensation.
  • Health, dental, and vision coverage for employees and dependents.
  • Wellness and commuter stipends for select roles, a 401(k) plan with a 2% company match for US employees, and flexible paid time off.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →