Site Reliability Engineer (Monetization)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
TL;DR
Site Reliability Engineer (Monetization): Own application-level infrastructure and reliability for a high-traffic commerce domain end to end, with an accent on Kubernetes deployments, SLO/SLI-driven observability, and production readiness. Focus on designing reliability practices early, running incident response and post-mortems, and evolving CI/CD and automation to keep distributed systems boringly reliable.
Location: Remote (Montreal)
Company
is a global commerce company providing tools and services to help game developers fund, distribute, market, and monetize their games.
What you will do
- Own application-level infrastructure for the Monetization domain, including Helm charts, Terraform configurations, Kubernetes deployments, runtime configuration, and service networking/integrations.
- Own domain observability by designing and implementing SLOs/SLIs, monitors, alerts, and dashboards using Datadog and OpenTelemetry-based tooling.
- Set up and evolve CI/CD pipelines for domain services (GitLab CI, GitHub Actions), including deploy and rollback automation.
- Perform capacity planning and performance tuning ahead of launches, sales events, and regional rollouts, including load testing and performance regression investigation.
- Run Production Readiness Reviews for new services and major changes; define and enforce what “production-ready” means for the domain.
- Support incident response with deep investigation, post-mortems, follow-up reliability improvements, and runbook maintenance; participate in SRE duty rotation.
Requirements
- 3+ years of SRE, DevOps, or platform engineering experience, including on-call/incident response, SLO/monitoring ownership, and production deploy pipeline/infrastructure work.
- Software development background: built and shipped backend services; comfortable reading application code during investigations and writing production-quality automation (e.g., Go, PHP).
- Hands-on Kubernetes experience (Helm, manifests, deploy strategies, debugging application-level performance and networking) on GKE or comparable managed Kubernetes.
- Strong observability practice: build monitors, dashboards, and SLOs/SLIs on modern tooling (Datadog preferred) and familiarity with OpenTelemetry.
- Infrastructure as Code exposure (Terraform/Terragrunt) and GCP experience (IAM, networking, managed services).
- Experience building/maintaining CI/CD pipelines (GitLab CI and/or GitHub Actions) and scripting/programming proficiency for automation (e.g., Python, Go, Bash).
Nice to have
- Kubernetes certifications.
- Google Cloud Platform certifications.
- HashiCorp certifications.
Culture & Benefits
- Comprehensive benefits program including medical, dental, and vision, plus PTO.
- Personalized career roadmap and support for professional development through training and educational opportunities.
- Supportive environment focused on physical, mental, and emotional well-being for employees and their families.
- Embedded role with daily collaboration with product development teams and shared reliability standards.
Hiring process
- Interviews and evaluations focused on SRE/production reliability experience, Kubernetes/observability, and incident response practices.
- Background check may include criminal history, employment verification, and education verification.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →