Назад
Company hidden
2 часа назад

Senior Site Reliability Engineer (Kubernetes)

Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Site Reliability Engineer (Kubernetes): Building and scaling reliable, resilient, and observable systems for high-traffic, customer-facing digital platforms with an accent on SLOs, observability, incident response, and automation. Focus on designing self-healing systems, improving deployment safety, analyzing performance bottlenecks, and engineering fault tolerance across distributed systems.

Location: Atlanta Support Center, United States; expected on-site presence of 80%

Company

Multi-brand restaurant company operating more than 33,300 Arby’s, Baskin-Robbins, Buffalo Wild Wings, Dunkin’, Jimmy John’s, and SONIC restaurants worldwide.

What you will do

  • Define and manage SLIs, SLOs, and error budgets for critical services.
  • Drive production readiness reviews, reliability requirements, capacity planning, failure mode analysis, and dependency risk assessments.
  • Design monitoring, alerting, logging, tracing, dashboards, and telemetry that accurately reflect service health.
  • Lead high-severity incident response, blameless postmortems, root cause analysis, and improvements to detection, response, and recovery.
  • Automate repetitive operational work by building self-healing systems, tooling, scripts, and safer CI/CD deployments.
  • Improve scalability and fault tolerance through load testing, performance analysis, resiliency patterns, and collaboration with engineering teams.

Requirements

  • 5+ years of experience in Site Reliability Engineering, Software Engineering, or Platform Engineering.
  • 2+ years of experience with Kubernetes and containerized workloads.
  • Four-year degree in Computer Science or a related field.
  • Strong programming or scripting skills in Python, Go, Java, or Node.
  • Experience operating against SLOs and error budgets, leading incident response, and performing root cause analysis.
  • Strong understanding of distributed systems, microservices architecture, cloud platforms, and observability strategies.

Nice to have

  • Experience with chaos engineering or resiliency testing.
  • Experience with high-volume, high-availability transactional systems.
  • Experience with AI-assisted observability or operational automation.
  • Contributions to internal SRE tooling, frameworks, or platforms.

Culture & Benefits

  • Engineering-driven reliability culture focused on reducing toil and preventing incidents.
  • Participation in an on-call rotation.
  • Mentorship and collaboration with engineering teams on SRE best practices.
  • Work supporting customer-facing digital platforms across a large restaurant brand portfolio.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →