5 дней назад
Senior Site Reliability Engineer (Kubernetes)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Site Reliability Engineer (Kubernetes): Building and operating reliable, scalable, and secure production infrastructure across cloud and on-premises environments with an accent on Kubernetes, GitOps, observability, and infrastructure as code. Focus on leading complex infrastructure initiatives, resolving database and production incidents, negotiating SLA-driven tradeoffs, and improving reliability, performance, and cost outcomes.
Location: Onsite in Jakarta, Indonesia
Company
develops digital payment and financial infrastructure products.
What you will do
- Own the availability, performance, scalability, and security of production systems across AWS/GCP and on-premises environments.
- Design and evolve Kubernetes deployment strategies for production workloads.
- Own production CI/CD and GitOps pipelines using ArgoCD or equivalent tools, along with Terraform-managed infrastructure.
- Build observability systems, diagnose database performance issues, and lead metrics-first incident investigations.
- Lead medium-to-large infrastructure initiatives, prioritize work by business impact, and negotiate technical tradeoffs with stakeholders to meet SLAs.
- Mentor engineers, promote operational best practices, and maintain documentation and incident processes.
Requirements
- At least 4 years of experience in SRE, DevOps, MLOps, or platform engineering, including senior-level ownership in a high-traffic environment.
- Deep expertise in at least one major cloud provider, preferably AWS, with the ability to quickly ramp up on another.
- Production experience with Kubernetes and strong Linux fundamentals.
- Experience with CI/CD, GitOps, ArgoCD or equivalent stacks, Terraform, and database performance analysis for MySQL or PostgreSQL.
- Experience with observability tools and standards such as Datadog and OpenTelemetry, plus on-call and incident handling during high-traffic events.
- Strong working English is required, along with clear verbal and written communication, documentation discipline, stakeholder negotiation, and the ability to own ambiguous scope.
Culture & Benefits
- Cross-functional collaboration with development, security, and product teams.
- Technical mentorship and guidance responsibilities across the engineering team.
- Ownership of end-to-end reliability, performance, security, and cost outcomes.
- Structured incident management and process improvement focused on high-traffic production systems.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →