Senior Platform SRE (AWS/Kubernetes)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
TL;DR
Senior Platform SRE (AWS/Kubernetes): Building and operating a reliability platform across AWS and on-premises HashiCorp Nomad with an accent on observability, SLOs, safe releases, and self-healing systems. Focus on designing chaos experiments, engineering automated remediation, improving distributed-system resilience, and establishing reliability standards for large-scale financial services platforms.
Location: City of London, United Kingdom; hybrid working with 3 days in the office
Company
is a FTSE 100 fintech operating trading and financial services platforms across five continents.
What you will do
- Build and own observability, SLO, error-budget, and burn-rate tracking capabilities using OpenTelemetry and distributed tracing.
- Establish 24/7 operational readiness through automated deployments, blue/green and canary releases, zero-downtime patching, and automated rollback.
- Engineer self-healing and traffic-management capabilities, including auto-remediation and error-budget-gated recovery.
- Design and execute controlled chaos experiments across AWS to identify and address reliability gaps.
- Build CI/CD and automation tools using software engineering practices, including version control, code reviews, and testing.
- Set reliability standards, support architecture and capacity reviews, facilitate blameless incident reviews, and mentor SREs and engineering teams.
Requirements
- Extensive production experience with OpenTelemetry, observability platforms such as Honeycomb, Datadog, Dynatrace, or Grafana, and direct instrumentation of Java or Python services.
- Proven ability to design meaningful SLIs and SLOs, configure multi-window burn-rate alerts, and manage error budgets with development teams.
- Experience building safe CI/CD pipelines with blue/green or canary releases, automated rollback, and DORA metrics.
- Kubernetes is required, together with cloud networking and infrastructure as code; Terraform is preferred. HashiCorp Nomad experience is advantageous.
- Production-quality Java and/or Python development experience and strong knowledge of distributed-system resilience patterns, including circuit breakers, bulkheads, idempotency, graceful degradation, and load shedding.
- Production on-call experience, blameless post-incident review facilitation, chaos engineering, strong troubleshooting, technical communication, and a bias toward automation.
Nice to have
- Experience in financial services, trading platforms, or other high-throughput, low-latency mission-critical environments.
- Experience with AWS FIS, Gremlin, PagerDuty, or ServiceNow.
- HashiCorp Nomad experience on a hybrid infrastructure estate.
Culture & Benefits
- Hybrid working model balancing office collaboration with flexibility.
- Tailored development programs, mentoring, career progression, and LinkedIn Learning access.
- Competitive salary with a flexible benefits package worth 12% of salary.
- Private medical cover, life insurance, gym membership contribution, and enhanced parental benefits.
- 28 days of total annual time off, including birthday and volunteering days, with the option to buy or sell holiday days.
- Employee-led inclusion networks, social clubs, and opportunities to participate in ESG initiatives.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →