17 часов назад
Site Reliability Engineer III (Platform)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer III (Platform) (AWS/Kubernetes): Designing and operating large-scale, distributed, fault-tolerant systems for Guidewire Cloud Platform and InsuranceSuite with an accent on automation, observability, and production reliability. Focus on building deployment and self-healing tools, optimizing microservice performance, and solving complex infrastructure issues through SLO tracking and incident learning.
Location: Kuala Lumpur, Malaysia
Company
provides cloud software, digital solutions, analytics, and AI-powered products for property and casualty insurance companies worldwide.
What you will do
- Collaborate with engineering teams, provide technical feedback, and contribute code that improves product functionality and resilience.
- Participate in on-call rotations and design tools supporting always-on, follow-the-sun operations for critical production systems.
- Automate deployments for core products and infrastructure using a robust automation framework.
- Monitor and optimize applications running on Cloud Platform to improve reliability and efficiency.
- Build observability tools, metrics, dashboards, and self-healing mechanisms while supporting SLO tracking and blameless postmortems.
- Identify infrastructure issues proactively and create system documentation and training materials.
Requirements
- Programming experience with Python or Go for internal tools, command-line interfaces, and APIs.
- Hands-on expertise with Docker, Helm, Kubernetes, EKS, CNI, Ingress networking, Kubernetes concepts, and the Operator pattern.
- Experience developing and testing complex Terraform modules and advanced AWS tooling with the AWS SDK.
- Knowledge of SSO, SAML, OAuth, observability tools such as Prometheus, OpenTelemetry, or Datadog, and CI/CD tools including TeamCity, GitHub Actions, or Jenkins.
- Production-at-scale support experience in heavily microservice-based environments, strong troubleshooting skills, and familiarity with Scrum or Kanban.
- Strong communication skills and the ability to apply emerging technologies, especially AI, to improve productivity and technical outcomes.
Nice to have
- Java and Spring Boot experience.
- Okta experience.
- Kubernetes or AWS certifications.
- Open-source contributions.
- Experience with KubeVela or Crossplane.
Culture & Benefits
- Participation in round-the-clock service operations through on-call and follow-the-sun practices.
- Focus on blameless postmortems, continuous learning, curiosity, innovation, and responsible use of AI.
- Collaborative and inclusive work environment centered on integrity, rationality, collegiality, innovation, and customer success.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
18 часов назад
Senior Site Reliability Engineer (Kubernetes)
13 часов назад
Staff Site Reliability Engineer (AI/ML)
241 000 - 270 000$
2 дня назад
Senior Site Reliability Engineer (AWS)
13 часов назад
Site Reliability Engineer Staff (Cloud Infrastructure)
18 часов назад
Site Reliability Engineer - Vice President (AI/RAG)
11 часов назад