8 дней назад
Site Reliability Engineer (Kubernetes/Cloud)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer (Kubernetes/Cloud): Building infrastructure automation and telemetry capabilities for a managed distributed database service across AWS, Azure, and Google Cloud with an accent on Kubernetes operations, cloud scalability, and customer-impact detection. Focus on debugging live production events, conducting postmortems and RCA analysis, and improving reliability through SLA-driven on-call operations.
Location: Portugal
Company
Cloud-native database provider delivering a distributed SQL database that unifies transactions and analytics for data-intensive applications.
What you will do
- Develop infrastructure automation for rollouts across multiple cloud providers.
- Optimize telemetry to identify customer-impacting events and provide data for debugging.
- Partner with engineering teams to improve service performance for cloud architectures.
- Debug live-site incidents and conduct follow-up postmortems and root-cause analysis.
- Participate in an SLA-driven on-call rotation, including after-hours, weekends, and rotating holidays.
Requirements
- Experience with infrastructure automation and production software troubleshooting.
- Knowledge of Kubernetes and the container ecosystem.
- Familiarity with at least one of AWS, Azure, or Google Cloud.
- Strong cross-group collaboration and communication skills.
- Bachelor’s degree in computer science or a related field.
Nice to have
- Experience with Python or Golang.
- Deep understanding of the query engine and backend infrastructure.
Culture & Benefits
- Work with a globally distributed engineering team.
- Help establish and scale site reliability practices across the company.
- Contribute to a cloud-focused managed database service operating across three major cloud providers.
- Work in a diverse and inclusive environment.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
9 дней назад
Site Reliability Engineer (AWS/Kubernetes)
13 дней назад
Sr. Site Reliability Engineer (Kubernetes/AWS)
Nscale
9 дней назад
Senior Site Reliability Engineer (AI Infrastructure Operations)
170 000 - 265 000$
Nscale
9 дней назад
Site Reliability Engineer (AI/GPU)
130 000 - 200 000$
9 дней назад
Senior Production Engineer (Fintech)
10 дней назад