Назад
Company hidden
15 часов назад

Senior Site Reliability Engineer (Data Platform)

Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Malaysia
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Site Reliability Engineer (Data Platform) (AWS): Building and operating reliable data platform services for large-scale processing, analytics, and streaming workloads with an accent on Kubernetes, infrastructure automation, observability, and cloud operations. Focus on designing self-service tooling, improving incident response, and implementing scalable, cost-efficient reliability patterns across distributed systems.

Location: Kuala Lumpur, Malaysia

Company

hirify.global provides a cloud platform for property and casualty insurers, combining digital, core, analytics, and AI capabilities.

What you will do

  • Design self-service automation and tooling in Go, Python, or scripting languages for deployment, operations, and troubleshooting.
  • Build and improve CI/CD pipelines with promotion gates and automated quality checks.
  • Provision and maintain cloud infrastructure for data and analytics workloads using Terraform and AWS CloudFormation.
  • Operate Kubernetes environments on AWS EKS, including deployment, scaling, lifecycle management, and containerized data services.
  • Develop observability for data platforms with metrics, traces, logs, dashboards, and actionable alerts using tools such as Datadog and ELK.
  • Improve reliability through capacity planning, cost-aware design, progressive delivery, chaos engineering, incident response, and collaboration with engineering, data, security, and SRE teams.

Requirements

  • 8+ years of experience in Site Reliability Engineering, DevOps, Production Engineering, or similar roles supporting distributed systems and data platforms.
  • Experience operating AWS cloud services in production, including on-call support.
  • Hands-on experience with Kafka, Hadoop, Spark, and Hive in public-cloud data platforms.
  • Proficiency in Java, Go, or Python, plus scripting for automation and integrations.
  • Strong experience with Kubernetes, Docker, CI/CD, Infrastructure as Code, microservices, REST or gRPC, and AWS data services.
  • Deep knowledge of distributed systems, networking, storage, operating systems, scalability, resilience, monitoring, logging, incident response, and troubleshooting.

Nice to have

  • Familiarity with Kubevela or Crossplane.
  • Kubernetes or AWS certifications.
  • Contributions to open-source projects.

Culture & Benefits

  • Collaborative work across engineering, data, operations, security, and SRE teams.
  • Culture emphasizing accountability, continuous learning, continuous improvement, collaboration, and psychological safety.
  • Agile development environment using Scrum and Kanban practices.
  • Focus on operational excellence, experimentation, and thoughtful adoption of emerging cloud, data, and SRE technologies.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →