3 дня назад
Site Reliability Engineer (SRE)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer (SRE): Ensuring resilience and performance across a school management platform with an accent on availability, scalability, observability, and capacity planning. Focus on monitoring platform performance, implementing SLOs, improving infrastructure through Terraform, and responding to incidents with runbooks, disaster recovery testing, and blameless postmortems.
Location: Remote within the United Kingdom
Company
develops MIS and school management tools used by more than 7,000 schools and trusts to improve staff workflows and turn data into actionable insights.
What you will do
- Monitor and analyse platform performance, availability, scalability, and capacity.
- Collaborate with engineering and platform teams to resolve bottlenecks and implement scalable solutions.
- Help define and review SLOs while improving monitoring, alerting, dashboards, and observability using tools such as Datadog and Prometheus.
- Design for high availability and resilience, create runbooks and playbooks, and test disaster recovery, backups, and high-availability plans.
- Respond to incidents, troubleshoot issues, minimise downtime, and participate in blameless postmortems.
- Work with support and other stakeholders to embed SRE practices and maintain service quality for customers.
Requirements
- Experience with performance monitoring and analysis and capacity planning.
- Scripting and automation skills, including Infrastructure as Code with Terraform.
- Understanding of relational databases and cloud versions such as AWS Aurora.
- Experience with messaging systems and distributed asynchronous workloads.
- Experience with nginx or similar technologies and familiarity with SRE processes.
- Understanding of DevOps principles, including the Three Ways and Five Ideals.
Nice to have
- Experience with other database technologies, cloud platforms, or enterprise solutions at scale.
- Familiarity with Kanban and Agile development processes.
- Experience with containerisation such as Docker.
- Knowledge of software practices including refactoring, clean code, domain-driven design, and test-driven development.
Culture & Benefits
- Flexible working arrangements.
- 32 days of holiday plus Bank Holidays.
- Wellbeing support, mindfulness initiatives, counselling, mental health resources, and a 24/7 virtual GP service.
- Life assurance worth three times annual salary, private dental insurance, and a pension scheme.
- Enhanced maternity, adoption, and paternity leave, plus return-to-work maternity coaching.
- Team and company events organised by social committees.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
3 дня назад
Platform Site Reliability Engineer (SRE)
9 дней назад
Site Reliability Engineer (SRE)
15 часов назад
DevOps/SRE Engineer
5 дней назад
Senior Infrastructure Engineer, SRE (AWS)
150 000 - 185 000$
8 дней назад
Senior Site Reliability Engineer (Observability)
160 000 - 200 000$
8 дней назад