Назад
Company hidden
5 дней назад

Senior Site Reliability Engineer (Kubernetes)

Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Indonesia
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Site Reliability Engineer (Kubernetes): Building and operating reliable, scalable, and secure production infrastructure across cloud and on-premises environments with an accent on Kubernetes, GitOps, observability, and infrastructure as code. Focus on leading complex infrastructure initiatives, resolving database and production incidents, negotiating SLA-driven tradeoffs, and improving reliability, performance, and cost outcomes.

Location: Onsite in Jakarta, Indonesia

Company

hirify.global develops digital payment and financial infrastructure products.

What you will do

  • Own the availability, performance, scalability, and security of production systems across AWS/GCP and on-premises environments.
  • Design and evolve Kubernetes deployment strategies for production workloads.
  • Own production CI/CD and GitOps pipelines using ArgoCD or equivalent tools, along with Terraform-managed infrastructure.
  • Build observability systems, diagnose database performance issues, and lead metrics-first incident investigations.
  • Lead medium-to-large infrastructure initiatives, prioritize work by business impact, and negotiate technical tradeoffs with stakeholders to meet SLAs.
  • Mentor engineers, promote operational best practices, and maintain documentation and incident processes.

Requirements

  • At least 4 years of experience in SRE, DevOps, MLOps, or platform engineering, including senior-level ownership in a high-traffic environment.
  • Deep expertise in at least one major cloud provider, preferably AWS, with the ability to quickly ramp up on another.
  • Production experience with Kubernetes and strong Linux fundamentals.
  • Experience with CI/CD, GitOps, ArgoCD or equivalent stacks, Terraform, and database performance analysis for MySQL or PostgreSQL.
  • Experience with observability tools and standards such as Datadog and OpenTelemetry, plus on-call and incident handling during high-traffic events.
  • Strong working English is required, along with clear verbal and written communication, documentation discipline, stakeholder negotiation, and the ability to own ambiguous scope.

Culture & Benefits

  • Cross-functional collaboration with development, security, and product teams.
  • Technical mentorship and guidance responsibilities across the engineering team.
  • Ownership of end-to-end reliability, performance, security, and cost outcomes.
  • Structured incident management and process improvement focused on high-traffic production systems.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →