1 день назад
Senior Site Reliability Specialist II
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Site Reliability Specialist II (Cloud Infrastructure/Kubernetes): Building reliable platform capabilities and improving the availability, scalability, performance, and resilience of critical production systems with an accent on cloud infrastructure, observability, automation, and distributed systems. Focus on leading cross-functional reliability initiatives, eliminating operational toil, improving incident response, and enabling engineering teams through self-service platforms and sound reliability practices.
Location: Remote in the United States
Company
provides Critical Event Management technology that combines intelligent automation and risk intelligence to help enterprises and government organizations manage critical events and protect people and operations.
What you will do
- Build platform capabilities and self-service tools that help engineering teams deliver reliable software safely and efficiently.
- Lead complex initiatives across cloud infrastructure, Kubernetes, observability, networking, automation, and developer platforms.
- Design solutions that improve platform availability, scalability, performance, resilience, recoverability, and operational readiness.
- Coach engineering teams on observability, incident response, disaster recovery, capacity planning, production readiness, SLOs, SLIs, and error budgets.
- Participate in on-call support, lead technical responses to high-severity incidents, and facilitate blameless post-incident reviews.
- Establish engineering standards, document best practices, review architectures, and drive corrective actions through completion.
Requirements
- Experience designing and operating complex production systems and cloud-native architectures.
- Experience with distributed systems, container platforms, Infrastructure as Code, automation, and CI/CD.
- Knowledge of observability, monitoring, logging, telemetry, incident response, and operational excellence.
- Understanding of reliability engineering principles, including SLOs, SLIs, capacity planning, and performance optimization.
- Ability to write software or automation in one or more modern programming languages.
- Linux and networking fundamentals.
Nice to have
- Experience working in regulated environments such as FedRAMP, DoD, IL4/IL5, SOC 2, or ISO 27001.
Culture & Benefits
- Collaborative technical leadership environment focused on ownership, continuous improvement, operational excellence, and customer outcomes.
- Opportunity to mentor engineers and influence architecture and engineering standards.
- Remote work arrangement for candidates in the United States.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
5 дней назад
Principal Production Engineer
160 200 - 425 000$
GitLab
5 дней назад
Site Reliability Engineer, Intermediate to Senior Staff (Kubernetes)
126 400 - 314 400$
5 дней назад
Site Reliability Engineer III (GCP)
4 дня назад
Platform Reliability Engineer (Kubernetes)
100 000 - 150 000$
6 дней назад
Senior Site Reliability Engineer (Fintech)
160 000 - 200 000$
7 дней назад