обновлено 3 дня назад
Senior Site Reliability Engineer (Kubernetes)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Site Reliability Engineer (Kubernetes): Running and improving production infrastructure for large distributed software applications with an accent on cloud reliability, automation, observability, and incident response. Focus on designing sustainable platform systems, optimizing performance, managing Kubernetes environments, and balancing delivery speed with service level objectives.
Location: Hybrid in London or Southampton, United Kingdom; 2 days in the office and 3 days remote each week
Company
develops AI, cloud, and digital software used by global businesses for customer experience, financial crime prevention, and public safety.
What you will do
- Monitor production availability and take a holistic view of system health.
- Build and operate software and systems for platform infrastructure and applications.
- Measure and optimize system performance, reliability, quality, and time-to-market.
- Provide operational support and engineering for multiple large distributed software applications.
- Gather and analyze operating system and application metrics for performance tuning and fault finding.
- Partner with development teams on system design, testing, release procedures, capacity planning, automation, and incident response.
Requirements
- 3–6 years of experience in systems engineering, automation, and reliability.
- Proficiency in at least one programming language such as Python, Go, Java, or C#, plus Bash or PowerShell scripting.
- Strong knowledge of AWS services, infrastructure as code with CloudFormation or Terraform, and CI/CD tools such as Jenkins, GitLab CI/CD, or CircleCI.
- Strong knowledge of Docker, Kubernetes, microservices architecture, and observability tools such as Prometheus, Grafana, ELK, or CloudWatch.
- Experience troubleshooting distributed systems, managing incidents, conducting blameless postmortems, and leading outage response and communication.
- Ability to work in the London or Southampton hybrid setup with two office days per week.
hirify.global-to-have"> to have
- Hands-on experience with large Kubernetes clusters and Kubernetes certification.
- Experience with the Grafana Observability Suite, including Loki, Mimir, and Tempo.
- Experience with Splunk, Datadog, PagerDuty, Rundeck, Ansible, Puppet, or Chef.
- AWS Certified DevOps Engineer, Google Cloud Professional DevOps Engineer, or equivalent certification.
Culture & Benefits
- -FLEX hybrid model with three remote workdays and two office days per week.
- Office days emphasize face-to-face meetings, teamwork, and collaborative problem-solving.
- Individual contributor role reporting to the Director of Network Operations.
- Equal opportunity employment environment.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
7 дней назад
Principal Site Reliability Engineer (AWS/Terraform)
107 000GBP
6 дней назад
Senior Site Reliability Engineer (AI)
7 дней назад
Senior DevOps Engineer
75 000 - 85 000$
10 дней назад
Lead Site Reliability Engineer (Kubernetes)
6 дней назад
Senior Site Reliability Engineer (AI)
75 000 - 85 000$
Anthropic
8 дней назад
Staff Software Engineer (AI Reliability Engineering)
325 000 - 390 000GBP