13 ΡΠ°ΡΠΎΠ² Π½Π°Π·Π°Π΄
Site Reliability Engineer (Azure)
ΠΡΡΡ & Π‘ΠΎΠΏΡΠΎΠ²ΠΎΠ΄
ΠΠ»Ρ ΠΌΡΡΡΠ° Ρ ΡΡΠΎΠΉ Π²Π°ΠΊΠ°Π½ΡΠΈΠ΅ΠΉ Π½ΡΠΆΠ΅Π½ Plus
ΠΠΏΠΈΡΠ°Π½ΠΈΠ΅ Π²Π°ΠΊΠ°Π½ΡΠΈΠΈ
Π’Π΅ΠΊΡΡ:
TL;DR
Site Reliability Engineer (Azure): Designing and operating observability, monitoring, and incident response solutions for hybrid and multi-cloud healthcare infrastructure with an accent on Azure, service reliability, and proactive alerting. Focus on automating self-healing systems, defining SLIs/SLOs/SLAs, improving incident resolution, and building resilient disaster recovery capabilities.
Location: Argentina β any location, remote
Company
develops virtual care and digital health services that support access to healthcare.
What you will do
- Design and maintain observability solutions across Azure and multi-cloud environments using monitoring, logging, tracing, dashboards, and alerting platforms.
- Define and standardize SLIs, SLOs, and SLAs to measure service health and customer experience.
- Build on-call runbooks and incident response playbooks, lead blameless retrospectives, and drive continuous improvement.
- Develop automation for self-healing systems, monitoring remediation, and incident mitigation.
- Contribute to disaster recovery and business continuity planning while improving resiliency, scalability, and reliability.
- Partner with security, network, and systems engineering teams and mentor engineers on observability and incident management practices.
Requirements
- 3β5 years of relevant experience or equivalent demonstrated through work experience, training, military experience, or education.
- Expertise with Azure VMs, AKS, Application Insights, and Azure Monitor.
- Hands-on experience with enterprise observability tools such as Datadog, Dynatrace, Elastic, Grafana, Prometheus, or LogicMonitor.
- Strong understanding of metrics, logs, traces, distributed-system monitoring, and automated alerting.
- Must be located in Argentina; the role is remote.
Nice to have
- Experience with Terraform, Bicep, Ansible, or similar infrastructure-as-code tools.
- Knowledge of Kubernetes observability in AKS environments.
- Exposure to AI-driven monitoring, anomaly detection, or predictive alerting.
- Scripting experience with Python, PowerShell, or similar languages.
- Experience with chaos engineering, healthcare IT, HIPAA, HITRUST, or other compliance frameworks.
Culture & Benefits
- Inclusive workplace focused on improving access to care and supporting diverse perspectives.
- Opportunities for professional growth, leadership, and meaningful impact.
- Benefits program designed around employees and their families.
- Innovative environment where fresh ideas and continuous improvement are valued.
Hiring process
- Identity and credential verification is conducted during the hiring process.
- Applicants complete live or video interviews.
- Fraud and misrepresentation screening is performed; falsified information leads to disqualification.
ΠΡΠ΄ΡΡΠ΅ ΠΎΡΡΠΎΡΠΎΠΆΠ½Ρ: Π΅ΡΠ»ΠΈ ΡΠ°Π±ΠΎΡΠΎΠ΄Π°ΡΠ΅Π»Ρ ΠΏΡΠΎΡΠΈΡ Π²ΠΎΠΉΡΠΈ Π² ΠΈΡ ΡΠΈΡΡΠ΅ΠΌΡ, ΠΈΡΠΏΠΎΠ»ΡΠ·ΡΡ iCloud/Google, ΠΏΡΠΈΡΠ»Π°ΡΡ ΠΊΠΎΠ΄/ΠΏΠ°ΡΠΎΠ»Ρ, Π·Π°ΠΏΡΡΡΠΈΡΡ ΠΊΠΎΠ΄/ΠΠ, Π½Π΅ Π΄Π΅Π»Π°ΠΉΡΠ΅ ΡΡΠΎΠ³ΠΎ - ΡΡΠΎ ΠΌΠΎΡΠ΅Π½Π½ΠΈΠΊΠΈ. ΠΠ±ΡΠ·Π°ΡΠ΅Π»ΡΠ½ΠΎ ΠΆΠΌΠΈΡΠ΅ "ΠΠΎΠΆΠ°Π»ΠΎΠ²Π°ΡΡΡΡ" ΠΈΠ»ΠΈ ΠΏΠΈΡΠΈΡΠ΅ Π² ΠΏΠΎΠ΄Π΄Π΅ΡΠΆΠΊΡ. ΠΠΎΠ΄ΡΠΎΠ±Π½Π΅Π΅ Π² Π³Π°ΠΉΠ΄Π΅ β
ΠΠΎΡ ΠΎΠΆΠΈΠ΅ Π²Π°ΠΊΠ°Π½ΡΠΈΠΈ
6 Π΄Π½Π΅ΠΉ Π½Π°Π·Π°Π΄
Site Reliability Engineer - DevSecOps Engineer (Cloud)
2 Π΄Π½Ρ Π½Π°Π·Π°Π΄
Site Reliability Engineer (Azure)
9 ΡΠ°ΡΠΎΠ² Π½Π°Π·Π°Π΄
Senior DevOps Engineer / Site Reliability Engineer (AI)
170Β 000 - 220Β 000$
11 ΡΠ°ΡΠΎΠ² Π½Π°Π·Π°Π΄
Lead Site Reliability Engineer (Azure)
Latitude
6 Π΄Π½Π΅ΠΉ Π½Π°Π·Π°Π΄
Senior Site Reliability Engineer (Kubernetes)
11 ΡΠ°ΡΠΎΠ² Π½Π°Π·Π°Π΄