Назад
Company hidden
5 дней назад

Staff Site Reliability Engineer (SaaS)

Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Staff Site Reliability Engineer (SaaS): Building reliable-by-default platforms, observability systems, change-safety tooling, and resilience automation for a global SaaS offering with an accent on distributed systems, Kubernetes, infrastructure as code, and SLO-driven operations. Focus on designing multi-region Azure services, leading incident learning, automating failure detection and response, and driving reliability adoption across product teams.

Location: Remote, United States. Standard shifts align with business hours in a 3×8h global rotation.

Total target compensation: $172,400–$441,500 USD depending on U.S. geographic zone.

Company

Data resilience and security software company developing a global SaaS offering for protecting organizational data and AI workloads.

What you will do

  • Build reliability features, reusable services, controllers, SDKs, and tooling adopted by product teams.
  • Define observability data models and implement metrics, logs, traces, SLI/SLO workflows, and error-budget policies.
  • Develop progressive delivery, automated rollback, release validation, fault injection, chaos experiments, and performance tooling.
  • Design and operate distributed, multi-region services initially running on Azure, with graceful degradation and strong operability.
  • Lead complex incidents, automate detection and response, and implement systemic fixes in code.
  • Mentor senior engineers and influence architecture through design reviews, ADRs, and cross-team initiatives.

Requirements

  • 8+ years of software engineering experience with cloud-based products and distributed systems at scale.
  • Production-grade backend development in at least one of C#, Java, Go, or TypeScript/Node.js.
  • Hands-on experience with Kubernetes, Terraform or Pulumi, and CI/CD systems such as GitHub Actions, GitLab, or ArgoCD.
  • Practical expertise in observability, metrics, tracing, logging, SLOs, and error budgets.
  • Ability to lead cross-team initiatives, influence architecture, and deliver measurable reliability outcomes.
  • Comfort with daytime on-call rotations and follow-the-sun coverage.

Nice to have

  • Experience building reliability platforms, progressive delivery systems, chaos tooling, or validation frameworks.
  • Multi-cloud experience or advanced Azure networking and traffic management.
  • Performance engineering at scale and security or compliance-aware delivery.

Culture & Benefits

  • Engineering-focused approach centered on preventing repeated incidents through code changes.
  • Data-driven reliability investment using SLIs, SLOs, and error budgets.
  • Blameless incident learning, paved roads, and self-service enablement for product teams.
  • Unlimited paid time off, paid holidays, volunteer hours, and parental leave.
  • Medical, dental, vision, mental health, retirement matching, and additional wellness and support programs.
  • Learning libraries, mentoring, workshops, and professional development events.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →