Назад
Company hidden
1 день назад

Senior Platform SRE (AWS/Kubernetes)

Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
UK
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/

TL;DR

Senior Platform SRE (AWS/Kubernetes): Building and operating a reliability platform across AWS and on-premises HashiCorp Nomad with an accent on observability, SLOs, safe releases, and self-healing systems. Focus on designing chaos experiments, engineering automated remediation, improving distributed-system resilience, and establishing reliability standards for large-scale financial services platforms.

Location: City of London, United Kingdom; hybrid working with 3 days in the office

Company

hirify.global is a FTSE 100 fintech operating trading and financial services platforms across five continents.

What you will do

  • Build and own observability, SLO, error-budget, and burn-rate tracking capabilities using OpenTelemetry and distributed tracing.
  • Establish 24/7 operational readiness through automated deployments, blue/green and canary releases, zero-downtime patching, and automated rollback.
  • Engineer self-healing and traffic-management capabilities, including auto-remediation and error-budget-gated recovery.
  • Design and execute controlled chaos experiments across AWS to identify and address reliability gaps.
  • Build CI/CD and automation tools using software engineering practices, including version control, code reviews, and testing.
  • Set reliability standards, support architecture and capacity reviews, facilitate blameless incident reviews, and mentor SREs and engineering teams.

Requirements

  • Extensive production experience with OpenTelemetry, observability platforms such as Honeycomb, Datadog, Dynatrace, or Grafana, and direct instrumentation of Java or Python services.
  • Proven ability to design meaningful SLIs and SLOs, configure multi-window burn-rate alerts, and manage error budgets with development teams.
  • Experience building safe CI/CD pipelines with blue/green or canary releases, automated rollback, and DORA metrics.
  • Kubernetes is required, together with cloud networking and infrastructure as code; Terraform is preferred. HashiCorp Nomad experience is advantageous.
  • Production-quality Java and/or Python development experience and strong knowledge of distributed-system resilience patterns, including circuit breakers, bulkheads, idempotency, graceful degradation, and load shedding.
  • Production on-call experience, blameless post-incident review facilitation, chaos engineering, strong troubleshooting, technical communication, and a bias toward automation.

Nice to have

  • Experience in financial services, trading platforms, or other high-throughput, low-latency mission-critical environments.
  • Experience with AWS FIS, Gremlin, PagerDuty, or ServiceNow.
  • HashiCorp Nomad experience on a hybrid infrastructure estate.

Culture & Benefits

  • Hybrid working model balancing office collaboration with flexibility.
  • Tailored development programs, mentoring, career progression, and LinkedIn Learning access.
  • Competitive salary with a flexible benefits package worth 12% of salary.
  • Private medical cover, life insurance, gym membership contribution, and enhanced parental benefits.
  • 28 days of total annual time off, including birthday and volunteering days, with the option to buy or sell holiday days.
  • Employee-led inclusion networks, social clubs, and opportunities to participate in ESG initiatives.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →