This vacancy is archived

View similar vacancies ↓
Company hidden
updated 13 days ago

Site Reliability Engineering Lead

Work type
fulltime
Grade
lead
English
c1
Country
Argentina

Job description

Text:
/
TL;DR
Site Reliability Engineering Lead (SRE): Driving the technical direction of production reliability, observability, and incident management for a global entertainment platform with an accent on reducing MTTD and MTTR. Focus on defining SLIs/SLOs, automating operational toil, and fostering a reliability-first engineering culture.

Location: Argentina

Company

World's leading tech platform for culture and live entertainment, democratizing access to experiences globally through a data-driven approach.

What you will do

  • Lead the SRE culture across the company, focusing on driving down MTTD and MTTR.
  • Define and continuously improve SLIs, SLOs, and Error Budgets to enhance service reliability.
  • Lead critical incident response, Root Cause Analysis (RCA), and implement corrective actions.
  • Improve monitoring, logging, tracing, and alerting to reduce operational toil.
  • Own operational processes including on-call practices, incident management, and runbooks.
  • Coach and develop the Tech Support Engineering team to foster a reliability-first mindset.

Requirements

  • Proven experience in Site Reliability Engineering, Production Engineering, or similar roles.
  • Experience leading engineering teams or high-impact technical initiatives.
  • Strong expertise in operating and troubleshooting large-scale distributed systems.
  • Hands-on experience with observability tools (Datadog, Grafana, Prometheus) and IaC (Terraform).
  • Proficiency in scripting (Python, Go, Bash) and solid software engineering experience.
  • English: Fluent proficiency required

Nice to have

  • Experience with Kubernetes and cloud platforms (AWS, GCP, or Azure).
  • Knowledge of OpenTelemetry or distributed tracing.
  • Experience with event-driven architectures such as Kafka.
  • Usage of AI-powered engineering tools (GitHub Copilot, ChatGPT) to improve productivity.

Hiring process

  • Talent Interview: Introduction to culture and background conversation.
  • Team Interview: Technical challenge and team-fit discussion.
  • Engineering Manager Interview: Deep dive into technical expertise and day-to-day operations.