17 часов назад
Staff Site Reliability Engineer (AWS/Kubernetes)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Staff Site Reliability Engineer (AWS/Kubernetes): Designing, building, and operating foundational infrastructure for Guidewire’s large-scale, multi-tenant SaaS platform with an accent on automation, reliability, scalability, and observability. Focus on developing self-healing systems, improving production resilience, leading incident response, and enabling follow-the-sun operations.
Location: Kuala Lumpur, Malaysia
Company
provides a cloud platform combining digital, core, analytics, and AI capabilities for property and casualty insurers worldwide.
What you will do
- Design, build, and operate reliable, scalable infrastructure for a multi-tenant SaaS platform.
- Automate deployment, provisioning, and operational workflows across AWS infrastructure and applications.
- Develop internal tools, services, and frameworks that reduce operational toil and improve engineering efficiency.
- Build observability systems using metrics, logging, tracing, and dashboards; define and track SLOs and reliability metrics.
- Lead or contribute to incident response, root cause analysis, blameless postmortems, and self-healing improvements.
- Collaborate with development teams on system design, production readiness, security, resilience, and scalability; mentor engineers and maintain runbooks.
Requirements
- Strong programming skills in Python or Go.
- Deep experience with AWS and production systems at scale.
- Hands-on expertise with Kubernetes, including EKS, Docker, Helm, CNI, Ingress networking, operators, RBAC, network policies, pod security standards, and secrets management.
- Experience with Infrastructure as Code using Terraform, Terragrunt, or similar tools.
- Solid understanding of Linux systems, networking, observability platforms, incident management, and microservices production support.
- Experience with CI/CD and GitOps tools such as GitHub Actions, TeamCity, Jenkins, FluxCD, or Bitbucket; ability to work in Scrum or Kanban environments.
Nice to have
- Java/Spring Boot, Kafka, SQS, Aurora, RDS, or Okta experience.
- AWS or Kubernetes certifications.
- Experience with large-scale SaaS platforms, KubeVela, Crossplane, or open-source contributions.
- Bachelor’s degree in Computer Science or a related field, or equivalent experience.
Culture & Benefits
- Work on a mission-critical global platform used by insurers in 40 countries.
- Collaborative engineering culture focused on integrity, rationality, collegiality, curiosity, and innovation.
- Opportunity to solve complex infrastructure problems at scale and shape an evolving cloud platform.
- Engineers are encouraged to use AI and emerging technologies to improve productivity and outcomes.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
18 часов назад
Senior Site Reliability Engineer (Kubernetes)
13 часов назад
Staff Site Reliability Engineer (AWS/Kubernetes)
13 часов назад
Staff Site Reliability Engineer (AI/ML)
241 000 - 270 000$
2 дня назад
Senior Site Reliability Engineer (AWS)
15 часов назад
Senior Site Reliability Engineer (GCP)
6 дней назад
Site Reliability Engineer (Kubernetes)
84 051 - 93 390GBP