обновлено 1 месяц назад
Lead, Site Reliability Engineer (Platform)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Lead, Site Reliability Engineer (Platform) (AWS/Kubernetes): Designing, building, and operating reliable infrastructure for Guidewire's multi-tenant SaaS platform with an accent on automation, observability, scalability, and security. Focus on building platform tools and frameworks, improving resilience, leading incident response, and enabling follow-the-sun production operations.
Location: Kuala Lumpur, Malaysia
Company
provides a cloud platform combining digital, core, analytics, and AI capabilities for property and casualty insurers worldwide.
What you will do
- Design, build, and operate reliable, scalable infrastructure for a multi-tenant SaaS platform.
- Automate deployment, provisioning, and operational workflows across cloud infrastructure and applications.
- Develop internal tools, services, and frameworks that reduce manual operational effort.
- Build observability systems, define SLOs, and improve platform resilience and self-healing capabilities.
- Lead or contribute to incident response, root cause analysis, and blameless postmortems.
- Collaborate with engineering teams, provide technical guidance, document operations, and mentor engineers.
Requirements
- Strong programming skills in Python or Go.
- Deep experience with AWS and production systems operating at scale.
- Hands-on expertise with Kubernetes, including EKS, Docker, Helm, CNI, Ingress, deployments, services, operators, and security primitives.
- Experience with Infrastructure as Code tools such as Terraform or Terragrunt, plus solid Linux and networking fundamentals.
- Experience with observability platforms such as Datadog, Prometheus, OpenTelemetry, or CloudWatch, as well as incident management in microservices environments.
- Working knowledge of SSO, SAML, OAuth, AWS IAM, CI/CD, GitOps, and secure cloud access patterns.
Nice to have
- Experience with Java/Spring Boot, Kafka, SQS, Aurora, or RDS.
- AWS or Kubernetes certifications.
- Experience with large-scale SaaS platforms, KubeVela, Crossplane, or open-source contributions.
- Bachelor's degree in Computer Science or a related field.
Culture & Benefits
- Work on a mission-critical global platform used by property and casualty insurers.
- Collaborate with distributed engineering teams in a high-impact culture focused on integrity, rationality, and collegiality.
- Shape the evolution of a cloud platform at significant scale.
- Use AI and emerging technologies to improve productivity and engineering outcomes.
- Participate in a follow-the-sun rotation supporting critical production systems.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →