15 часов назад
Senior Site Reliability Engineer (Data Platform)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Site Reliability Engineer (Data Platform) (AWS): Building and operating reliable data platform services for large-scale processing, analytics, and streaming workloads with an accent on Kubernetes, infrastructure automation, observability, and cloud operations. Focus on designing self-service tooling, improving incident response, and implementing scalable, cost-efficient reliability patterns across distributed systems.
Location: Kuala Lumpur, Malaysia
Company
provides a cloud platform for property and casualty insurers, combining digital, core, analytics, and AI capabilities.
What you will do
- Design self-service automation and tooling in Go, Python, or scripting languages for deployment, operations, and troubleshooting.
- Build and improve CI/CD pipelines with promotion gates and automated quality checks.
- Provision and maintain cloud infrastructure for data and analytics workloads using Terraform and AWS CloudFormation.
- Operate Kubernetes environments on AWS EKS, including deployment, scaling, lifecycle management, and containerized data services.
- Develop observability for data platforms with metrics, traces, logs, dashboards, and actionable alerts using tools such as Datadog and ELK.
- Improve reliability through capacity planning, cost-aware design, progressive delivery, chaos engineering, incident response, and collaboration with engineering, data, security, and SRE teams.
Requirements
- 8+ years of experience in Site Reliability Engineering, DevOps, Production Engineering, or similar roles supporting distributed systems and data platforms.
- Experience operating AWS cloud services in production, including on-call support.
- Hands-on experience with Kafka, Hadoop, Spark, and Hive in public-cloud data platforms.
- Proficiency in Java, Go, or Python, plus scripting for automation and integrations.
- Strong experience with Kubernetes, Docker, CI/CD, Infrastructure as Code, microservices, REST or gRPC, and AWS data services.
- Deep knowledge of distributed systems, networking, storage, operating systems, scalability, resilience, monitoring, logging, incident response, and troubleshooting.
Nice to have
- Familiarity with Kubevela or Crossplane.
- Kubernetes or AWS certifications.
- Contributions to open-source projects.
Culture & Benefits
- Collaborative work across engineering, data, operations, security, and SRE teams.
- Culture emphasizing accountability, continuous learning, continuous improvement, collaboration, and psychological safety.
- Agile development environment using Scrum and Kanban practices.
- Focus on operational excellence, experimentation, and thoughtful adoption of emerging cloud, data, and SRE technologies.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
2 дня назад
Senior Site Reliability Engineer (AWS)
13 часов назад
Site Reliability Engineer Staff (Cloud Infrastructure)
18 часов назад
Senior Site Reliability Engineer (Kubernetes)
10 часов назад
Senior Infrastructure Engineer (Kubernetes)
5 дней назад
Senior Site Reliability Engineer (Kubernetes)
93 700 - 138 700$
15 часов назад