1 Π΄Π΅Π½Ρ Π½Π°Π·Π°Π΄
Site Reliability Engineer (Kubernetes)
ΠΡΡΡ & Π‘ΠΎΠΏΡΠΎΠ²ΠΎΠ΄
ΠΠ»Ρ ΠΌΡΡΡΠ° Ρ ΡΡΠΎΠΉ Π²Π°ΠΊΠ°Π½ΡΠΈΠ΅ΠΉ Π½ΡΠΆΠ΅Π½ Plus
ΠΠΏΠΈΡΠ°Π½ΠΈΠ΅ Π²Π°ΠΊΠ°Π½ΡΠΈΠΈ
Π’Π΅ΠΊΡΡ:
TL;DR
Site Reliability Engineer (Kubernetes) (Commerce Infrastructure): Building and operating application-level infrastructure and reliability systems for a high-traffic commerce domain with an accent on Kubernetes, observability, CI/CD, and production readiness. Focus on designing SLOs and SLIs, automating deployments and rollbacks, planning capacity for major launches, and investigating complex production incidents.
Location: Baku, Azerbaijan; hybrid workplace
Company
is a global commerce company providing tools and services that help video game developers fund, distribute, market, and monetize their games.
What you will do
- Own application-level infrastructure, including Helm charts, Terraform configurations, Kubernetes deployments, runtime configuration, networking, and integrations.
- Design and operate observability for critical services using SLOs, SLIs, monitors, alerts, dashboards, Datadog, and OpenTelemetry-based tooling.
- Build and evolve CI/CD pipelines with GitLab CI and GitHub Actions, including deployment and rollback automation.
- Perform capacity planning, load testing, performance tuning, and regression investigation for launches, sales events, and regional rollouts.
- Lead production readiness reviews, incident investigations, post-mortems, runbook maintenance, and reliability improvements.
- Partner with product engineering teams on planning, architecture reviews, reliability roadmaps, and operational standards.
Requirements
- 3+ years of SRE, DevOps, or platform engineering experience with production infrastructure, on-call duties, incident response, monitoring, and deployment pipelines.
- Software development experience building and shipping backend services, plus production-quality automation in Go, PHP, Python, Bash, or a comparable language.
- Hands-on Kubernetes experience, including Helm, manifests, deployment strategies, and debugging performance and networking issues.
- Experience with observability platforms and SLO/SLI implementation; Datadog is preferred, while Prometheus and Grafana are relevant.
- Experience with Terraform or Terragrunt, GCP infrastructure, IAM, networking, managed services, and GitLab CI or GitHub Actions.
- Practical incident response experience and a background in payments, fintech, e-commerce, gaming, or other high-traffic transactional systems.
Nice to have
- Kubernetes, Google Cloud Platform, or HashiCorp certifications.
Culture & Benefits
- Supportive and collaborative working environment.
- Medical, dental, and vision coverage.
- Paid time off and benefits supporting employee and family well-being.
- Personalized career roadmap, training, and educational opportunities.
- Inclusive culture focused on creativity, collaboration, and the gaming industry.
ΠΡΠ΄ΡΡΠ΅ ΠΎΡΡΠΎΡΠΎΠΆΠ½Ρ: Π΅ΡΠ»ΠΈ ΡΠ°Π±ΠΎΡΠΎΠ΄Π°ΡΠ΅Π»Ρ ΠΏΡΠΎΡΠΈΡ Π²ΠΎΠΉΡΠΈ Π² ΠΈΡ ΡΠΈΡΡΠ΅ΠΌΡ, ΠΈΡΠΏΠΎΠ»ΡΠ·ΡΡ iCloud/Google, ΠΏΡΠΈΡΠ»Π°ΡΡ ΠΊΠΎΠ΄/ΠΏΠ°ΡΠΎΠ»Ρ, Π·Π°ΠΏΡΡΡΠΈΡΡ ΠΊΠΎΠ΄/ΠΠ, Π½Π΅ Π΄Π΅Π»Π°ΠΉΡΠ΅ ΡΡΠΎΠ³ΠΎ - ΡΡΠΎ ΠΌΠΎΡΠ΅Π½Π½ΠΈΠΊΠΈ. ΠΠ±ΡΠ·Π°ΡΠ΅Π»ΡΠ½ΠΎ ΠΆΠΌΠΈΡΠ΅ "ΠΠΎΠΆΠ°Π»ΠΎΠ²Π°ΡΡΡΡ" ΠΈΠ»ΠΈ ΠΏΠΈΡΠΈΡΠ΅ Π² ΠΏΠΎΠ΄Π΄Π΅ΡΠΆΠΊΡ. ΠΠΎΠ΄ΡΠΎΠ±Π½Π΅Π΅ Π² Π³Π°ΠΉΠ΄Π΅ β
ΠΠΎΡ ΠΎΠΆΠΈΠ΅ Π²Π°ΠΊΠ°Π½ΡΠΈΠΈ
2 Π΄Π½Ρ Π½Π°Π·Π°Π΄
Site Reliability Engineer (Aerospace)
2 Π΄Π½Ρ Π½Π°Π·Π°Π΄
Site Reliability Engineer (TypeScript)
180Β 000 - 220Β 000$
1 Π΄Π΅Π½Ρ Π½Π°Π·Π°Π΄
Senior Site Reliability Engineer (Kubernetes)
128Β 500 - 190Β 000$
2 Π΄Π½Ρ Π½Π°Π·Π°Π΄
Site Reliability Engineer (Network)
1 Π΄Π΅Π½Ρ Π½Π°Π·Π°Π΄
Site Reliability Engineer (AWS)
2 Π΄Π½Ρ Π½Π°Π·Π°Π΄
Principal Site Reliability Engineer (AWS/Kubernetes)
163Β 620 - 212Β 710$