Senior Cloud Site Reliability Engineer
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
TL;DR
Senior Cloud Site Reliability Engineer (Cloud/AWS): Running production environments through monitoring and holistic system health, building platform infrastructure and applications, and improving reliability, quality, and time-to-market with an accent on distributed systems troubleshooting, automation, and incident management. Focus on performance tuning, capacity planning, and partnering with development teams to strengthen testing and release procedures.
Location: United Kingdom (London, Southampton)
Company
develops software products used by global businesses to deliver customer experiences, fight financial crime, and ensure public safety.
What you will do
- Run production by monitoring availability and maintaining a holistic view of system health.
- Build software and systems to manage platform infrastructure and applications.
- Improve reliability, quality, and time-to-market across a suite of distributed software applications.
- Measure and optimize system performance using metrics from operating systems and applications.
- Provide primary operational support and engineering, including incident response and blameless postmortems.
- Automate and uplift services while balancing delivery speed with service level objectives.
Requirements
- 3–6 years of experience in a similar role focused on systems engineering, automation, and reliability.
- Proficiency in at least one programming language (e.g., Python, Go, Java, C#) and experience with scripting (e.g., Bash, PowerShell).
- Deep understanding of cloud platforms and reliability constraints (e.g., AWS services such as EC2, ECS, Lambda, DynamoDB).
- Experience with infrastructure as code (e.g., CloudFormation, Terraform).
- Strong CI/CD knowledge and experience with tools such as Jenkins, GitLab CI/CD, or CircleCI.
- Experience with containerization and orchestration (e.g., Docker, Kubernetes) plus monitoring/observability tools (e.g., Prometheus, Grafana, ELK, CloudWatch).
hirify.global-to-have"> to have
- Hands-on experience with large Kubernetes clusters and relevant certifications.
- Experience with Grafana Observability Suite (Loki, Mimir, Tempo).
- Experience with monitoring/automation tools such as Splunk, Datadog, PagerDuty, Rundeck.
- Familiarity with configuration management tools like Ansible, Puppet, or Chef.
- AWS or Google Cloud DevOps certifications (or equivalent).
Culture & Benefits
- -FLEX hybrid model: 2 days in the office and 3 days remote each week.
- Office days emphasize face-to-face meetings and collaborative problem-solving.
- Role is an Individual Contributor position reporting into the Director, Network Operations.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →