23 часа назад
Site Reliability Engineer (SRE)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer (SRE) (Grafana/Prometheus): Maintaining the health, performance, and observability of multiple UK e-commerce client platforms with an accent on incident response, monitoring, and operational documentation. Focus on improving dashboards and alerts, performing root cause analysis, automating recurring operational work, and supporting reliable handoffs across time zones.
Location: Latin America, preferably aligned with PT or ET working hours. Candidates in Santiago, Chile may work remotely or from the Santiago office; candidates outside Santiago will be fully remote.
Company
delivers technology and managed services for international clients, including UK e-commerce platforms.
What you will do
- Monitor platform health across multiple client environments using Grafana, Prometheus, and comparable observability tools.
- Respond to, triage, classify, and escalate incidents using established runbooks and workflows.
- Maintain observability dashboards, alerts, SLI, and SLO tracking.
- Support the Service Desk with technical triage, resolution workflows, and clear communication with technical and non-technical stakeholders.
- Conduct root cause analysis, participate in post-incident reviews, and maintain postmortems and operational documentation.
- Identify recurring issues and implement automation or process improvements to reduce operational toil.
Requirements
- Strong English proficiency for written and verbal communication.
- 2–3 years of experience in SRE, platform operations, or a technical Service Desk role.
- Experience with monitoring and observability tools such as Grafana and Prometheus.
- Knowledge of incident management, including triage, escalation, and postmortems, plus experience supporting e-commerce platforms.
- Scripting experience with shell and/or Python; operational familiarity with Docker and Kubernetes.
- Experience with Agile environments, ticketing tools such as Jira, testing, code review, and independent early-morning shifts.
Nice to have
- Experience with Google Cloud Platform, AWS, Azure, or another major cloud provider.
- Familiarity with CI/CD pipelines such as GitHub Actions or GitLab CI.
- Basic experience with Terraform and Infrastructure as Code.
- Knowledge of AIOps concepts and relevant SRE or cloud certifications.
Culture & Benefits
- Work across a range of client engagements with international brands.
- Inclusive and safe working environment.
- Training budgets and mentorship focused on Agentic AI and professional skills.
- Generous vacation policy.
- Weekend on-call rotation with an alternating one-weekend-on, one-weekend-off schedule.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
6 дней назад
Senior Site Reliability Engineer (Observability)
4 дня назад
Site Reliability Engineer (SRE)
100 000 - 180 000$
3 дня назад
Site Reliability Engineer (NOC)
4 дня назад
DevOps Engineer (SRE, Kubernetes)
1 день назад
Engineer III, Site Reliability (SRE)
4 дня назад
Reliability Engineer (SRE)
75 000 - 95 000$