5 дней назад
Staff Site Reliability Engineer (SaaS)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Staff Site Reliability Engineer (SaaS): Building reliable-by-default platforms, observability systems, change-safety tooling, and resilience automation for a global SaaS offering with an accent on distributed systems, Kubernetes, infrastructure as code, and SLO-driven operations. Focus on designing multi-region Azure services, leading incident learning, automating failure detection and response, and driving reliability adoption across product teams.
Location: Remote, United States. Standard shifts align with business hours in a 3×8h global rotation.
Total target compensation: $172,400–$441,500 USD depending on U.S. geographic zone.
Company
Data resilience and security software company developing a global SaaS offering for protecting organizational data and AI workloads.
What you will do
- Build reliability features, reusable services, controllers, SDKs, and tooling adopted by product teams.
- Define observability data models and implement metrics, logs, traces, SLI/SLO workflows, and error-budget policies.
- Develop progressive delivery, automated rollback, release validation, fault injection, chaos experiments, and performance tooling.
- Design and operate distributed, multi-region services initially running on Azure, with graceful degradation and strong operability.
- Lead complex incidents, automate detection and response, and implement systemic fixes in code.
- Mentor senior engineers and influence architecture through design reviews, ADRs, and cross-team initiatives.
Requirements
- 8+ years of software engineering experience with cloud-based products and distributed systems at scale.
- Production-grade backend development in at least one of C#, Java, Go, or TypeScript/Node.js.
- Hands-on experience with Kubernetes, Terraform or Pulumi, and CI/CD systems such as GitHub Actions, GitLab, or ArgoCD.
- Practical expertise in observability, metrics, tracing, logging, SLOs, and error budgets.
- Ability to lead cross-team initiatives, influence architecture, and deliver measurable reliability outcomes.
- Comfort with daytime on-call rotations and follow-the-sun coverage.
Nice to have
- Experience building reliability platforms, progressive delivery systems, chaos tooling, or validation frameworks.
- Multi-cloud experience or advanced Azure networking and traffic management.
- Performance engineering at scale and security or compliance-aware delivery.
Culture & Benefits
- Engineering-focused approach centered on preventing repeated incidents through code changes.
- Data-driven reliability investment using SLIs, SLOs, and error budgets.
- Blameless incident learning, paved roads, and self-service enablement for product teams.
- Unlimited paid time off, paid holidays, volunteer hours, and parental leave.
- Medical, dental, vision, mental health, retirement matching, and additional wellness and support programs.
- Learning libraries, mentoring, workshops, and professional development events.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
6 дней назад
Site Reliability Engineering Team Lead (Principal SRE, Automotive AI)
132 000 - 211 400$
5 дней назад
SRE (Kubernetes/DevOps)
Baseten
8 дней назад
Site Reliability Engineer (AI)
165 000 - 330 000$
6 дней назад
Senior/Staff Site Reliability Engineer
77 000 - 94 300€
Chess.com
8 дней назад
Site Reliability Engineer (Gaming)
Baseten
8 дней назад
Site Reliability Engineer (AI)
165 000 - 330 000$