Ssr Monitoring and Observability Analyst (SRE)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
TL;DR
Ssr Monitoring and Observability Analyst (SRE): Designing and maintaining proactive monitoring and alerting systems for global clients' IT infrastructure with an accent on SLI/SLO management and high availability. Focus on building end-to-end observability solutions, reducing MTTR through Root Cause Analysis, and automating incident response using AIOps.
Location: Remote (Must be based in Argentina, Uruguay, Mexico, or Colombia)
Company
designs and delivers scalable digital solutions for global businesses with a strong technical foundation and a product-driven mindset.
What you will do
- Define and execute the company’s observability strategy following SRE and DevOps best practices.
- Configure business-impact-driven alert thresholds (SLIs/SLOs) to reduce notification noise.
- Develop and maintain real-time dashboards using Grafana and Kibana for operational visibility.
- Implement monitoring automation from agent deployment to basic/intermediate AIOps incident response.
- Administer and optimize monitoring platforms to ensure stability and manage infrastructure costs.
- Author service maps, monitoring runbooks, and troubleshooting procedures to streamline incident response.
Requirements
- Must be based in Argentina, Uruguay, Mexico, or Colombia.
- 3+ years of experience in Monitoring, IT Operations, SRE, or Systems Administration.
- Advanced expertise with Prometheus, Grafana, ELK Stack, New Relic, or Datadog.
- Hands-on experience monitoring Cloud environments (AWS, Azure, or GCP) and containerized workloads (Docker, Kubernetes).
- Knowledge of log aggregation (Fluentd, Logstash, Loki) and Distributed Tracing (Jaeger, Zipkin, OpenTelemetry).
- Practical proficiency in Python or Bash for custom checker creation and automation.
Nice to have
- Official Cloud Certifications (AWS, Azure, or GCP).
- Tooling Certifications (Datadog, Dynatrace, Elastic, Prometheus).
- SRE or DevOps certifications and foundational knowledge.
- Solid grasp of core networking concepts (TCP/IP, DNS, Load Balancing).
Culture & Benefits
- 100% remote work with a long-term commitment and high autonomy.
- Strategic, high-visibility role within a modern engineering culture.
- Opportunity to work with a collaborative international team and strong technical leadership.
- Clear path to professional growth and leadership positions within the company.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →