обновлено 10 дней назад
Senior Site Reliability Engineer (Kubernetes)
147 600 - 221 400$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Site Reliability Engineer (Kubernetes/Cloud Infrastructure): Building and operating reliable, scalable cloud infrastructure and observability systems for a large enterprise platform with an accent on Kubernetes, SLIs/SLOs, incident response, and AWS/Azure networking. Focus on diagnosing production failures, improving distributed-system reliability, automating operational work, and defining scalability and availability requirements.
Location: Remote within the United States
Salary: $147,600–$221,400 USD annually in Zone 1; $137,900–$206,900 USD annually in Zone 2, depending on the US location. Total compensation also includes an annual bonus and equity.
Company
provides cloud-based software that helps service businesses operate more efficiently.
What you will do
- Participate in an on-call rotation, diagnose production issues, and perform root-cause analysis and remediation.
- Design and maintain observability dashboards and alerting based on SLIs and SLOs.
- Operate and improve the Kubernetes-based compute platform running most of the infrastructure.
- Work across AWS and Azure cloud networking and infrastructure to support reliable, scalable systems.
- Partner with product engineering teams on architecture, infrastructure, non-functional requirements, and reliability best practices.
- Build automation, maintain runbooks and documentation, and contribute to CI/CD pipelines.
Requirements
- Strong hands-on knowledge of Kubernetes and practical experience applying SRE principles, including SLIs, SLOs, and error budgets.
- Solid AWS or Azure cloud engineering and networking fundamentals, including subnetting and IP addressing.
- Deep experience with an observability stack such as OpenTelemetry, Prometheus, Grafana, Datadog, or Elasticsearch.
- Strong CI/CD experience; GitHub Actions is preferred, with TeamCity, Azure DevOps, or GitLab CI also accepted.
- Strong programming skills for building web applications, ideally with .NET and ASP.NET, or with Python and Flask/FastAPI or Java and Spring.
- 8–10+ years of relevant hands-on experience, plus experience with distributed systems and production troubleshooting.
Nice to have
- Database experience.
Culture & Benefits
- Flexible time off, flextime, autonomous work support, and learning and development opportunities.
- Comprehensive onboarding and leadership training.
- Company-paid medical, dental, and vision insurance, with FSA, HSA, 401(k) match, and telehealth options.
- Parental leave, fertility, surrogacy, and adoption support, plus maternity and breast milk shipping programs.
- Pet insurance, legal advisory services, financial planning tools, recognition programs, and peer-nominated awards.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
9 дней назад
Senior Site Reliability Engineer (Cloud-Native Infrastructure)
142 800 - 178 500$
12 дней назад
Sr Site Reliability Engineer (AWS)
95 000 - 135 000$
12 дней назад
Site Reliability Engineer (AWS)
110 000 - 140 000$
12 дней назад
Senior Site Reliability Engineer (Kubernetes)
8 дней назад
DevOps & SRE Engineer (Kubernetes)
100 000 - 150 000$
12 дней назад