Назад
Company hidden
5 дней назад

Site Reliability Engineer III (Data Platform)

Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Malaysia
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Site Reliability Engineer III (Data Platform) (AWS/Kubernetes): Operating and evolving large-scale data platform services for AI and analytics products with an accent on reliability, automation, and production operations. Focus on managing Kafka, Hadoop, Spark, and Hive, building infrastructure automation with Go or Python, and improving incident response, scaling, and progressive delivery on AWS.

Location: Kuala Lumpur, Malaysia

Company

hirify.global provides a cloud platform for property and casualty insurers, combining digital, core, analytics, and AI capabilities.

What you will do

  • Operate and enhance production data platform services to maintain high availability, performance, and cost efficiency.
  • Manage and optimize Kafka, Hadoop, Spark, and Hive platforms on AWS, including configuration, tuning, and daily operations.
  • Define SLOs, error budgets, capacity plans, and scaling strategies for data platform components.
  • Participate in on-call rotations, triage alerts with PagerDuty, resolve production incidents, and contribute to blameless post-incident reviews.
  • Build automation and troubleshooting tools with Go, Python, or scripting, and support CI/CD pipelines with TeamCity or GitHub Actions.
  • Maintain infrastructure with Terraform or AWS CloudFormation, operate Kubernetes environments on AWS EKS, and implement blue/green and canary deployments.

Requirements

  • 4–6 years of experience in Site Reliability Engineering, DevOps, Production Engineering, or similar roles supporting distributed systems or data platforms.
  • Experience deploying and operating services on AWS or Azure, including on-call support.
  • Expertise with Kafka, Hadoop, Spark, or Hive, plus proficiency in Go, Python, or Bash.
  • Experience with CI/CD, Infrastructure as Code, Kubernetes, Docker, observability, and logging tools.
  • Understanding of distributed systems, networking, storage, operating systems, and incident response.
  • Bachelor’s or master’s degree in Computer Science, Computer Engineering, Mathematics, or a related technical field, or equivalent practical experience.

Nice to have

  • Knowledge of Java and Spring Boot.
  • Familiarity with Kubevela or Crossplane.
  • Kubernetes or AWS certifications.
  • Contributions to open source projects.

Culture & Benefits

  • Collaboration with product engineering, data platform, and security teams.
  • Agile practices including Scrum and Kanban.
  • Blameless incident reviews focused on reliability improvements, runbooks, and automation.
  • Work supporting AI, cloud, analytics, and data platform adoption.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →