Назад
4 часа назад

Senior Site Reliability Engineer (AI)

240 000 - 356 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/

TL;DR

Senior Site Reliability Engineer (Kubernetes/AI): Improving the reliability, scalability, and operational maturity of compute provisioning and infrastructure orchestration systems with an accent on Kubernetes, infrastructure automation, and observability. Focus on designing fault-isolation mechanisms, automating detection of configuration drift, and leading production incident response for AI workloads.

Location: Hybrid: Must be based in San Francisco, San Jose, or Bellevue (presence required 4 days per week)

Salary: $240,000 – $356,000

Company

Lambda is a leader in AI cloud infrastructure providing compute power for AI researchers and enterprises.

What you will do

  • Operate and scale critical platform services across physical data centers.
  • Improve reliability of compute provisioning, instance lifecycle, and regional orchestration.
  • Build monitoring, alerting, and tracing for service health and provisioning latency.
  • Define SLIs, SLOs, error budgets, and operational readiness standards.
  • Automate detection and remediation of configuration drift and failed workflows.
  • Lead production incident response, postmortems, and durable corrective actions.

Requirements

  • 7+ years of experience in SRE, infrastructure, or distributed systems.
  • Deep production experience operating Kubernetes architecture and networking.
  • Proficiency with Terraform or similar infrastructure-as-code tools.
  • Experience building CI/CD or GitOps workflows using Argo CD, Flux, Helm, or Kustomize.
  • Experience with observability platforms such as OpenTelemetry, Prometheus, Grafana, or Datadog.
  • Ability to develop production-quality tooling in Go or Python.

Nice to have

  • Experience with AI infrastructure, GPU platforms, or high-performance computing.
  • Knowledge of Kubernetes controllers, operators, CRDs, or scheduler extensions.
  • Experience with chaos engineering, fault injection, or automated remediation.
  • Familiarity with SOC 2, ISO 27001, or similar compliance frameworks.

Culture & Benefits

  • Competitive cash and equity compensation.
  • Comprehensive health, dental, and vision coverage for employees and dependents.
  • 401k Plan with 2% company match for USA employees.
  • Flexible paid time off plan.
  • Wellness and commuter stipends for select roles.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →