4 дня назад
Site Reliability Engineer III, GWCP (SaaS)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer III, GWCP (SaaS): Building and operating reliable, scalable infrastructure for a multi-tenant SaaS platform with an accent on automation, observability, security, and production readiness. Focus on Kubernetes and AWS operations, self-healing systems, incident response, and enabling follow-the-sun reliability across distributed teams.
Location: Curitiba, Brazil
Company
provides a cloud platform for property and casualty insurers, combining core insurance operations with digital, data, analytics, and AI products.
What you will do
- Design, build, and operate reliable, scalable infrastructure for a multi-tenant SaaS platform.
- Automate deployment, provisioning, and operational workflows across cloud infrastructure and applications.
- Develop internal tools, services, and frameworks that reduce manual operational effort.
- Build observability systems, define SLOs, and improve reliability metrics across production services.
- Lead or contribute to incident response, root cause analysis, blameless postmortems, and resilience improvements.
- Collaborate with development teams on system design, security, production readiness, documentation, and technical enablement.
Requirements
- Strong programming skills in Python or Go; Java/Spring Boot is an advantage.
- Deep experience with AWS and production systems operating at scale.
- Hands-on experience with Kubernetes, including EKS, Docker, Helm, CNI, Ingress, and Kubernetes primitives.
- Experience with Infrastructure as Code using Terraform, Terragrunt, or similar tools.
- Knowledge of Linux systems, networking, observability, incident management, CI/CD, and GitOps.
- Working knowledge of SSO, SAML, OAuth, identity providers, and security and compliance standards.
Nice to have
- Experience with Kafka, SQS, Aurora, RDS, Okta, KubeVela, or Crossplane.
- Bachelor’s degree in Computer Science or a related field, or equivalent experience.
- Experience supporting large-scale SaaS platforms.
- AWS or Kubernetes certifications and open-source contributions.
Culture & Benefits
- Work on a mission-critical cloud platform used by more than 540 insurers in 40 countries.
- Participate in a 24x7 follow-the-sun on-call rotation supporting critical production systems.
- Collaborate with distributed engineering teams in a culture focused on integrity, innovation, learning, and collegiality.
- Use AI and data-driven insights to improve engineering productivity and outcomes.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
4 дня назад
Site Reliability Engineer (AWS)
5 дней назад
Senior Site Reliability Engineer (GCP)
2 дня назад
Sr. Site Reliability Engineer (AWS/CI/CD)
4 дня назад
Site Reliability Engineer II, GovCloud
103 000 - 155 000$
5 дней назад
Site Reliability Engineer Staff (Cloud Infrastructure)
6 дней назад