обновлено 22 часа назад
Site Reliability Engineer - Vice President (Kubernetes)
130 000 - 160 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer - Vice President (Kubernetes) (AWS/Observability): Designing and operating reliable, scalable Kubernetes-based services and observability systems for iCapital’s platform with an accent on SLOs, SLIs, monitoring as code, and incident response. Focus on standardizing telemetry, automating remediation, leading high-severity incidents, and driving systemic reliability improvements across distributed production systems.
Location: Salt Lake City, Utah, United States; office work Monday–Thursday with remote work available on Friday
Salary: $130,000–$160,000 base salary annually, depending on level
Company
provides a platform for its client base and operates a Site Reliability Engineering function focused on consistent, reliable service delivery.
What you will do
- Define and iterate service level objectives and indicators that reflect customer and business expectations.
- Standardize monitoring and alerting through monitors as code, preferably with Terraform, including severity, ownership, and runbook quality gates.
- Develop observability standards across metrics, logs, and traces, including OpenTelemetry instrumentation and dependency mapping.
- Define reliability and operability standards for Kubernetes services, including scaling, resource constraints, rollout safety, dashboards, and alerts.
- Automate incident workflows, runbooks, remediation, and other toil-reduction initiatives.
- Serve as Incident Commander, lead postmortems, participate in on-call rotations, and drive measurable reliability improvements.
Requirements
- 7+ years of experience in SRE or related roles, demonstrating technical seniority across multiple services and teams.
- Strong production experience with AWS and Kubernetes.
- Experience defining SLOs and SLIs and applying them to operational and engineering decisions.
- Strong Infrastructure as Code skills, preferably Terraform, with experience building reusable automation and configuration standards.
- Experience with data stores and managed services such as Postgres, MongoDB, or DynamoDB, including distributed-system failure modes.
- Experience with at least two observability stacks, strong incident response and debugging skills, and clear written and verbal communication.
Culture & Benefits
- Hybrid schedule with office collaboration Monday–Thursday and remote flexibility on Friday.
- Salary, equity for all full-time employees, and an annual performance bonus.
- Employer-matched retirement plan and subsidized healthcare.
- Employer-paid dental, vision, telemedicine, and virtual mental health counseling.
- Parental leave and unlimited paid time off.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
2 дня назад
Site Reliability Engineer (AWS)
120 000 - 185 000$
Okta
5 дней назад
Manager, Site Reliability Engineering (AWS/Kubernetes)
204 000 - 306 000$
Okta
6 дней назад
Manager, Site Reliability Engineering (AWS/Kubernetes)
204 000 - 306 000$
2 дня назад
Staff Site Reliability Engineer (AI/ML)
112 500 - 187 500$
7 дней назад
Site Reliability Engineer II (AWS)
100 000 - 110 000$
Okta
6 дней назад
Senior Manager, Site Reliability Engineering (Federal)
207 000 - 284 900$