7 дней назад
Lead Site Reliability Engineer (Kubernetes)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Lead Site Reliability Engineer (Kubernetes): Building and maintaining reliable, scalable, and high-performing production infrastructure with an accent on Kubernetes operations, automation, and observability. Focus on incident response, defining SLOs and SLIs, optimizing capacity and performance, and reducing operational toil through automation.
Location: London, United Kingdom; hybrid work model required
Company
provides financial data, analytics, and software solutions for investment professionals and financial organizations worldwide.
What you will do
- Monitor, maintain, and improve the reliability and availability of production systems.
- Respond to incidents, perform post-mortems, and implement measures to prevent recurrence.
- Define and track Service Level Objectives and Service Level Indicators.
- Collaborate with development teams to build reliability into services from the beginning.
- Design automation that reduces operational toil and improves efficiency.
- Support capacity planning, performance optimization, documentation, runbooks, and the on-call rotation.
Requirements
- Hands-on experience deploying, managing, troubleshooting, and administering Kubernetes workloads and clusters.
- Strong knowledge of Pods, Deployments, Services, ConfigMaps, Ingress, Kubernetes networking, storage, and security.
- Experience with Helm for application packaging and deployment.
- Knowledge of cloud platforms, CI/CD tooling, monitoring and observability, infrastructure as code, configuration management, and scripting or programming.
- Bachelor’s degree in computer science or a relevant field.
- Fluent English required in verbal and written communication; willingness to work in a hybrid model required.
Nice to have
- Experience contributing to open-source projects.
- Familiarity with the SRE principles described in the Google SRE handbook.
- Previous experience in DevOps or Platform Engineering.
Culture & Benefits
- Blameless incident response culture.
- Focus on continuous learning and improvement.
- Collaboration across technical and non-technical teams.
- Participation in an on-call rotation supporting critical systems.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →