обновлено 1 день назад
Site Reliability Technical Lead
80 000 - 90 000GBP
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Technical Lead (AWS/Terraform/Kubernetes): Designing and operating robust, scalable, and secure cloud platforms serving approximately 1,000 requests per second with an accent on reliability, observability, automation, and distributed systems. Focus on defining architecture, leading incident analysis, building CI/CD and infrastructure automation, and mentoring engineers toward production-ready services.
Location: Remote from the United Kingdom
Salary: £80,000–£90,000 per year
Company
develops MIS and school management tools used by more than 7,000 schools and trusts.
What you will do
- Define and guide system architecture, balancing scalability, maintainability, security, and delivery speed.
- Improve platform reliability, performance, efficiency, observability, and adherence to Service Level Objectives.
- Lead incident response and Root Cause Analysis, while optimizing incident management practices.
- Drive automation, infrastructure improvements, production readiness, and operational toil reduction.
- Lead technical estimation, feasibility assessments, release planning, and post-release reviews.
- Mentor engineers and collaborate with Product Managers, Engineering Managers, and technical stakeholders.
Requirements
- Extensive experience in SRE, DevOps, or Platform Engineering for complex, scalable systems.
- Advanced expertise with AWS, distributed cloud architectures, distributed systems, microservices, and resilience patterns.
- Experience operating platforms handling approximately 1,000 requests per second.
- Advanced Terraform and configuration management skills, plus proficiency in Python, Go, or a similar language.
- Experience with monitoring and observability platforms such as DataDog or Prometheus, incident management, Docker, Kubernetes, ECS, and CI/CD pipelines.
- Demonstrated ability to mentor and support the growth of other engineers.
Nice to have
- Experience with chaos engineering and reliability testing.
- Knowledge of security best practices and compliance frameworks.
- Background in agile and lean methodologies such as Scrum or Kanban.
- Open-source or SRE community contributions.
Culture & Benefits
- Flexible working arrangements and a dedicated wellbeing team.
- 32 days of holiday plus Bank Holidays.
- Life assurance at three times annual salary, private dental insurance, pension, and enhanced parental leave.
- Virtual GP, mental health support, counselling, health checks, and financial wellbeing services.
- Professional development training budget and one paid volunteering day each year.
- Social committees and team, office, and company-wide events.
Hiring process
- Phone screen.
- First-stage interview.
- Second-stage interview.
Visa sponsorship is not available.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
2 дня назад
Senior Site Reliability Engineer (Kubernetes)
2 дня назад
Cloud Infrastructure Engineer (AWS)
60 000 - 65 000GBP
3 дня назад
Site Reliability Engineer (AWS/Kubernetes)
90 000 - 120 000GBP
2 дня назад
Principal Cloud Engineer (AI)
85 000 - 125 000GBP
5 дней назад
Lead DevOps Engineer (AWS)
10 часов назад