2 часа назад
Senior Site Reliability Engineer (Kubernetes)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Site Reliability Engineer (Kubernetes): Building and scaling reliable, resilient, and observable systems for high-traffic, customer-facing digital platforms with an accent on SLOs, observability, incident response, and automation. Focus on designing self-healing systems, improving deployment safety, analyzing performance bottlenecks, and engineering fault tolerance across distributed systems.
Location: Atlanta Support Center, United States; expected on-site presence of 80%
Company
Multi-brand restaurant company operating more than 33,300 Arby’s, Baskin-Robbins, Buffalo Wild Wings, Dunkin’, Jimmy John’s, and SONIC restaurants worldwide.
What you will do
- Define and manage SLIs, SLOs, and error budgets for critical services.
- Drive production readiness reviews, reliability requirements, capacity planning, failure mode analysis, and dependency risk assessments.
- Design monitoring, alerting, logging, tracing, dashboards, and telemetry that accurately reflect service health.
- Lead high-severity incident response, blameless postmortems, root cause analysis, and improvements to detection, response, and recovery.
- Automate repetitive operational work by building self-healing systems, tooling, scripts, and safer CI/CD deployments.
- Improve scalability and fault tolerance through load testing, performance analysis, resiliency patterns, and collaboration with engineering teams.
Requirements
- 5+ years of experience in Site Reliability Engineering, Software Engineering, or Platform Engineering.
- 2+ years of experience with Kubernetes and containerized workloads.
- Four-year degree in Computer Science or a related field.
- Strong programming or scripting skills in Python, Go, Java, or Node.
- Experience operating against SLOs and error budgets, leading incident response, and performing root cause analysis.
- Strong understanding of distributed systems, microservices architecture, cloud platforms, and observability strategies.
Nice to have
- Experience with chaos engineering or resiliency testing.
- Experience with high-volume, high-availability transactional systems.
- Experience with AI-assisted observability or operational automation.
- Contributions to internal SRE tooling, frameworks, or platforms.
Culture & Benefits
- Engineering-driven reliability culture focused on reducing toil and preventing incidents.
- Participation in an on-call rotation.
- Mentorship and collaboration with engineering teams on SRE best practices.
- Work supporting customer-facing digital platforms across a large restaurant brand portfolio.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
7 дней назад
Staff Site Reliability Engineer (AI/Blockchain)
195 000 - 257 500$
3 дня назад
Systems Reliability Engineer (Kubernetes/Cloud)
100 000 - 150 000$
4 дня назад
Site Reliability Engineer (SRE)
100 000 - 180 000$
Okta
5 дней назад
Staff Site Reliability Engineer (Kubernetes)
194 000 - 267 000$
24 часа назад
Senior Site Reliability Specialist II
1 день назад