11 дней назад
Lead Site Reliability Engineer (Kubernetes)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Lead Site Reliability Engineer (Kubernetes): Building and maintaining reliable, scalable, and highly available production infrastructure with an accent on Kubernetes administration, automation, observability, and incident response. Focus on defining SLOs and SLIs, reducing operational toil, optimizing performance, and designing systems that remain resilient under critical workloads.
Location: São Paulo, Brazil; hybrid work required
Company
provides financial data, analytics, and software solutions for investment professionals and financial institutions worldwide.
What you will do
- Monitor, maintain, and improve the reliability and availability of production systems.
- Respond to incidents, conduct post-mortems, and prevent recurring failures.
- Define and track Service Level Objectives and Service Level Indicators.
- Collaborate with development teams to build reliability into services.
- Design automation that reduces operational toil and improves efficiency.
- Contribute to on-call support, capacity planning, performance optimization, documentation, and runbooks.
Requirements
- Hands-on experience deploying, managing, troubleshooting, and administering Kubernetes workloads and clusters.
- Strong knowledge of Pods, Deployments, Services, ConfigMaps, Ingress, Kubernetes networking, storage, and security.
- Experience with Helm for application packaging and deployment.
- Bachelor's degree in computer science or a relevant field.
- Fluent English required, both verbal and written.
- Willingness to work in a hybrid model in São Paulo.
Nice to have
- Experience with cloud platforms such as AWS, GCP, or Azure.
- Experience with CI/CD, monitoring and observability, infrastructure as code, configuration management, or scripting.
- Open-source contributions or familiarity with the Google SRE principles.
- Previous DevOps or Platform Engineering experience.
Culture & Benefits
- Participation in an on-call rotation supporting critical systems.
- Blameless incident response culture and continuous learning.
- Collaboration across technical and non-technical teams.
- Focus on automation, continuous improvement, and methodical problem-solving.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →