2 часа назад
Senior Platform SRE (AWS)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Platform SRE (AWS): Building and owning a reliability platform across HashiCorp Nomad and AWS with an accent on observability, SLOs, automated releases, and self-healing capabilities. Focus on designing chaos experiments, implementing reliability patterns in Java or Python, and leading incident learning for high-throughput production systems.
Location: Kraków, Poland; hybrid model with 3 days in the office
Company
is a FTSE 100 fintech operating across five continents, serving over 1.3 million customers and processing billions of dollars in transactions.
What you will do
- Build and own the reliability platform across on-premises HashiCorp Nomad and AWS.
- Implement OpenTelemetry-based monitoring, distributed tracing, SLOs, error budgets, and burn-rate tracking.
- Establish operational readiness through automated deployments, blue/green releases, zero-downtime patching, and automated rollback.
- Engineer self-healing capabilities, traffic rerouting, CI/CD automation, and reliability-focused tooling.
- Design and run controlled chaos experiments across the AWS estate.
- Define SRE standards, mentor engineers, facilitate blameless post-incident reviews, and drive remediation actions to closure.
Requirements
- 6+ years of experience across observability, SLOs, CI/CD, container orchestration, software engineering, distributed systems, incident management, and chaos engineering.
- Hands-on OpenTelemetry experience and production use of Honeycomb, Datadog, Dynatrace, or Grafana.
- Production-quality Java and/or Python development experience, including implementation of reliability patterns in application codebases.
- Kubernetes is required; experience with EKS, AKS, or GKE, cloud networking, and infrastructure as code, preferably Terraform.
- Experience with on-call production support, blameless PIR facilitation, contributing-factor analysis, and remediation tracking.
- Strong communication, systems thinking, automation, troubleshooting, and cross-functional collaboration skills.
Nice to have
- HashiCorp Nomad experience.
- Experience with AWS FIS, Gremlin, or equivalent chaos engineering tools.
- PagerDuty and ServiceNow familiarity.
- Experience in financial services, trading platforms, or other mission-critical, high-throughput environments.
Culture & Benefits
- Hybrid working model designed to balance office collaboration and connection.
- Tailored development programs, mentoring opportunities, and clear career progression.
- Committees, sports clubs, and social clubs for professional and community engagement.
- Additional time off for volunteering and community work.
- Culture focused on leadership, commercial impact, client focus, sustainable delivery, and accountability.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
Affirm
6 дней назад
Senior Site Reliability Engineer (SRE & Platform Reliability)
308 000 - 428 000PLN
1 день назад
Senior Site Reliability Engineer (Kubernetes)
19 часов назад
Principal Site Reliability Engineer (AWS)
220 000 - 280 000PLN
16 часов назад
Staff Software Engineer (Infrastructure, Cloud)
3 дня назад
Staff Site Reliability Engineer (AI)
20 часов назад