Senior Site Reliability Engineer (AWS)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
TL;DR
Senior Site Reliability Engineer (AWS/Kubernetes): Maintaining reliability, performance, and availability of Develocity instances with an accent on automation, observability, and incident response. Focus on building SRE practices from the ground up, automating cloud infrastructure on AWS, and optimizing SaaS operations.
Location: Remote from anywhere in Europe in the GMT timezone
Company
is an AI-native company building Develocity, a toolchain observability and intelligence platform used by leading global software organizations.
What you will do
- Operate and maintain all Develocity instances and supporting infrastructure services.
- Participate in a follow-the-sun on-call rotation, owning incident response and troubleshooting across the stack.
- Drive automation for application deployment, upgrades, monitoring, self-healing, and recovery.
- Build and maintain comprehensive observability using logging, metrics, tracing, and alerting.
- Collaborate with engineering teams to build reliability directly into features from inception.
- Own disaster recovery planning, backups, and business continuity strategies.
Requirements
- 5+ years in SRE, DevOps, or equivalent role operating production services at scale.
- Must be based in Europe in the GMT timezone.
- Strong production experience with Kubernetes and AWS (EKS, RDS, S3, EC2).
- Proficiency with Infrastructure as Code (Terraform) and observability tools (Prometheus, Grafana).
- Scripting proficiency in Python and Bash for automation.
- Strong written and verbal English communication.
Nice to have
- Experience operating large-scale SaaS platforms.
- Proficiency in JVM languages such as Java or Kotlin.
- Experience establishing SRE practices in new or growing teams.
- Proven track record in disaster recovery planning and execution.
- Familiarity with Develocity.
Culture & Benefits
- Founding member role in a new SRE team with significant ownership and influence.
- Remote-first environment emphasizing asynchronous communication and written documentation.
- Competitive salary and equity grants.
- Culture that prioritizes automation over manual "heroics".
- Annual company offsites and team meetings for in-person collaboration.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →