Назад
Company hidden
7 дней назад

Senior Site Reliability Engineer (AWS)

Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US/Switzerland
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Site Reliability Engineer (AWS) (SaaS/Serverless): Building the reliability practice for an API-first, composable banking platform on AWS with an accent on observability, incident response, and resilient serverless operations. Focus on defining SLOs, automating detection and remediation, improving GitHub Actions delivery pipelines, and designing disaster recovery for regulated client-facing environments.

Location: Fort Lauderdale, Florida, United States; hybrid and flexible working, with regular travel to Switzerland required.

Company

hirify.global provides wealth management technology and services, including a SaaS, API-first, composable banking platform for financial institutions.

What you will do

  • Define and evolve SLIs, SLOs, and error budgets with product teams to guide reliability decisions.
  • Design observability for distributed serverless systems across metrics, logs, and traces.
  • Build incident response practices covering on-call operations, escalation, blameless post-mortems, and remediation.
  • Develop reliability automation that detects and resolves issues before they affect clients.
  • Improve GitHub Actions CI/CD, deployment automation, progressive delivery, capacity planning, and resilient system design.
  • Contribute to disaster recovery, operational readiness, regulatory compliance, and mentoring across engineering teams.

Requirements

  • 5+ years of experience in Site Reliability Engineering, DevOps, or production operations for distributed cloud systems.
  • Substantial hands-on AWS experience and practical knowledge of SLIs, SLOs, and error budgets.
  • Strong incident response experience, including on-call work, pressure triage, and post-mortems.
  • Clear understanding of observability for distributed systems.
  • Strong automation and scripting skills; Rust, TypeScript, or Python experience is relevant.
  • Excellent collaboration and communication skills with a pragmatic approach to speed and stability.

Nice to have

  • AWS serverless and event-driven architecture experience, including idempotency, retries, dead-letter queues, and failure handling.
  • Experience taking platforms from pre-launch to production and defining operational readiness criteria.
  • Experience with AI agents for platform health monitoring and AI-assisted engineering tools.
  • GitHub Actions, Terraform, OpenTofu, progressive delivery, and disaster recovery experience.
  • Experience with regulated B2B SaaS environments, SOC 2, PCI DSS, GDPR, multi-cloud operations, or relevant AWS certifications.

Culture & Benefits

  • Hybrid and flexible working is available for most employees.
  • Work with a global organization operating across 10 countries and partnering with financial institutions worldwide.
  • Collaborative environment focused on shared operational ownership and inclusive teamwork.
  • Regular collaboration with the Zurich SRE team and platform engineering teams.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →