Назад
Company hidden
1 дСнь назад

Site Reliability Engineer (Kubernetes)

Π€ΠΎΡ€ΠΌΠ°Ρ‚ Ρ€Π°Π±ΠΎΡ‚Ρ‹
hybrid
Π’ΠΈΠΏ Ρ€Π°Π±ΠΎΡ‚Ρ‹
fulltime
Π“Ρ€Π΅ΠΉΠ΄
middle
Английский
b2
Π‘Ρ‚Ρ€Π°Π½Π°
US/Azerbaijan
Вакансия ΠΈΠ· списка Hirify.GlobalВакансия ΠΈΠ· Hirify RU Global, списка ΠΊΠΎΠΌΠΏΠ°Π½ΠΈΠΉ с восточно-СвропСйскими корнями
Для мэтча ΠΈ ΠΎΡ‚ΠΊΠ»ΠΈΠΊΠ° Π½ΡƒΠΆΠ΅Π½ Plus

ΠœΡΡ‚Ρ‡ & Π‘ΠΎΠΏΡ€ΠΎΠ²ΠΎΠ΄

Для мэтча с этой вакансиСй Π½ΡƒΠΆΠ΅Π½ Plus

ОписаниС вакансии

ВСкст:
/
TL;DR
Site Reliability Engineer (Kubernetes) (Commerce Infrastructure): Building and operating application-level infrastructure and reliability systems for a high-traffic commerce domain with an accent on Kubernetes, observability, CI/CD, and production readiness. Focus on designing SLOs and SLIs, automating deployments and rollbacks, planning capacity for major launches, and investigating complex production incidents.

Location: Baku, Azerbaijan; hybrid workplace

Company

hirify.global is a global commerce company providing tools and services that help video game developers fund, distribute, market, and monetize their games.

What you will do

  • Own application-level infrastructure, including Helm charts, Terraform configurations, Kubernetes deployments, runtime configuration, networking, and integrations.
  • Design and operate observability for critical services using SLOs, SLIs, monitors, alerts, dashboards, Datadog, and OpenTelemetry-based tooling.
  • Build and evolve CI/CD pipelines with GitLab CI and GitHub Actions, including deployment and rollback automation.
  • Perform capacity planning, load testing, performance tuning, and regression investigation for launches, sales events, and regional rollouts.
  • Lead production readiness reviews, incident investigations, post-mortems, runbook maintenance, and reliability improvements.
  • Partner with product engineering teams on planning, architecture reviews, reliability roadmaps, and operational standards.

Requirements

  • 3+ years of SRE, DevOps, or platform engineering experience with production infrastructure, on-call duties, incident response, monitoring, and deployment pipelines.
  • Software development experience building and shipping backend services, plus production-quality automation in Go, PHP, Python, Bash, or a comparable language.
  • Hands-on Kubernetes experience, including Helm, manifests, deployment strategies, and debugging performance and networking issues.
  • Experience with observability platforms and SLO/SLI implementation; Datadog is preferred, while Prometheus and Grafana are relevant.
  • Experience with Terraform or Terragrunt, GCP infrastructure, IAM, networking, managed services, and GitLab CI or GitHub Actions.
  • Practical incident response experience and a background in payments, fintech, e-commerce, gaming, or other high-traffic transactional systems.

Nice to have

  • Kubernetes, Google Cloud Platform, or HashiCorp certifications.

Culture & Benefits

  • Supportive and collaborative working environment.
  • Medical, dental, and vision coverage.
  • Paid time off and benefits supporting employee and family well-being.
  • Personalized career roadmap, training, and educational opportunities.
  • Inclusive culture focused on creativity, collaboration, and the gaming industry.

Π‘ΡƒΠ΄ΡŒΡ‚Π΅ остороТны: Ссли Ρ€Π°Π±ΠΎΡ‚ΠΎΠ΄Π°Ρ‚Π΅Π»ΡŒ просит Π²ΠΎΠΉΡ‚ΠΈ Π² ΠΈΡ… систСму, ΠΈΡΠΏΠΎΠ»ΡŒΠ·ΡƒΡ iCloud/Google, ΠΏΡ€ΠΈΡΠ»Π°Ρ‚ΡŒ ΠΊΠΎΠ΄/ΠΏΠ°Ρ€ΠΎΠ»ΡŒ, Π·Π°ΠΏΡƒΡΡ‚ΠΈΡ‚ΡŒ ΠΊΠΎΠ΄/ПО, Π½Π΅ Π΄Π΅Π»Π°ΠΉΡ‚Π΅ этого - это мошСнники. ΠžΠ±ΡΠ·Π°Ρ‚Π΅Π»ΡŒΠ½ΠΎ ΠΆΠΌΠΈΡ‚Π΅ "ΠŸΠΎΠΆΠ°Π»ΠΎΠ²Π°Ρ‚ΡŒΡΡ" ΠΈΠ»ΠΈ ΠΏΠΈΡˆΠΈΡ‚Π΅ Π² ΠΏΠΎΠ΄Π΄Π΅Ρ€ΠΆΠΊΡƒ. ΠŸΠΎΠ΄Ρ€ΠΎΠ±Π½Π΅Π΅ Π² Π³Π°ΠΉΠ΄Π΅ β†’