Назад
Company hidden
обновлено 22 часа назад

Site Reliability Engineer - Vice President (Kubernetes)

130 000 - 160 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
director
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Site Reliability Engineer - Vice President (Kubernetes) (AWS/Observability): Designing and operating reliable, scalable Kubernetes-based services and observability systems for iCapital’s platform with an accent on SLOs, SLIs, monitoring as code, and incident response. Focus on standardizing telemetry, automating remediation, leading high-severity incidents, and driving systemic reliability improvements across distributed production systems.

Location: Salt Lake City, Utah, United States; office work Monday–Thursday with remote work available on Friday

Salary: $130,000–$160,000 base salary annually, depending on level

Company

hirify.global provides a platform for its client base and operates a Site Reliability Engineering function focused on consistent, reliable service delivery.

What you will do

  • Define and iterate service level objectives and indicators that reflect customer and business expectations.
  • Standardize monitoring and alerting through monitors as code, preferably with Terraform, including severity, ownership, and runbook quality gates.
  • Develop observability standards across metrics, logs, and traces, including OpenTelemetry instrumentation and dependency mapping.
  • Define reliability and operability standards for Kubernetes services, including scaling, resource constraints, rollout safety, dashboards, and alerts.
  • Automate incident workflows, runbooks, remediation, and other toil-reduction initiatives.
  • Serve as Incident Commander, lead postmortems, participate in on-call rotations, and drive measurable reliability improvements.

Requirements

  • 7+ years of experience in SRE or related roles, demonstrating technical seniority across multiple services and teams.
  • Strong production experience with AWS and Kubernetes.
  • Experience defining SLOs and SLIs and applying them to operational and engineering decisions.
  • Strong Infrastructure as Code skills, preferably Terraform, with experience building reusable automation and configuration standards.
  • Experience with data stores and managed services such as Postgres, MongoDB, or DynamoDB, including distributed-system failure modes.
  • Experience with at least two observability stacks, strong incident response and debugging skills, and clear written and verbal communication.

Culture & Benefits

  • Hybrid schedule with office collaboration Monday–Thursday and remote flexibility on Friday.
  • Salary, equity for all full-time employees, and an annual performance bonus.
  • Employer-matched retirement plan and subsidized healthcare.
  • Employer-paid dental, vision, telemedicine, and virtual mental health counseling.
  • Parental leave and unlimited paid time off.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →