Назад
Company hidden
обновлено 1 месяц назад

Lead, Site Reliability Engineer (Platform)

Формат работы
onsite
Тип работы
fulltime
Грейд
lead
Английский
b2
Страна
Malaysia
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Lead, Site Reliability Engineer (Platform) (AWS/Kubernetes): Designing, building, and operating reliable infrastructure for Guidewire's multi-tenant SaaS platform with an accent on automation, observability, scalability, and security. Focus on building platform tools and frameworks, improving resilience, leading incident response, and enabling follow-the-sun production operations.

Location: Kuala Lumpur, Malaysia

Company

hirify.global provides a cloud platform combining digital, core, analytics, and AI capabilities for property and casualty insurers worldwide.

What you will do

  • Design, build, and operate reliable, scalable infrastructure for a multi-tenant SaaS platform.
  • Automate deployment, provisioning, and operational workflows across cloud infrastructure and applications.
  • Develop internal tools, services, and frameworks that reduce manual operational effort.
  • Build observability systems, define SLOs, and improve platform resilience and self-healing capabilities.
  • Lead or contribute to incident response, root cause analysis, and blameless postmortems.
  • Collaborate with engineering teams, provide technical guidance, document operations, and mentor engineers.

Requirements

  • Strong programming skills in Python or Go.
  • Deep experience with AWS and production systems operating at scale.
  • Hands-on expertise with Kubernetes, including EKS, Docker, Helm, CNI, Ingress, deployments, services, operators, and security primitives.
  • Experience with Infrastructure as Code tools such as Terraform or Terragrunt, plus solid Linux and networking fundamentals.
  • Experience with observability platforms such as Datadog, Prometheus, OpenTelemetry, or CloudWatch, as well as incident management in microservices environments.
  • Working knowledge of SSO, SAML, OAuth, AWS IAM, CI/CD, GitOps, and secure cloud access patterns.

Nice to have

  • Experience with Java/Spring Boot, Kafka, SQS, Aurora, or RDS.
  • AWS or Kubernetes certifications.
  • Experience with large-scale SaaS platforms, KubeVela, Crossplane, or open-source contributions.
  • Bachelor's degree in Computer Science or a related field.

Culture & Benefits

  • Work on a mission-critical global platform used by property and casualty insurers.
  • Collaborate with distributed engineering teams in a high-impact culture focused on integrity, rationality, and collegiality.
  • Shape the evolution of a cloud platform at significant scale.
  • Use AI and emerging technologies to improve productivity and engineering outcomes.
  • Participate in a follow-the-sun rotation supporting critical production systems.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →