7 часов назад
Senior Site Reliability Engineer (Azure)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Site Reliability Engineer (Azure/Observability): Operating, scaling, and improving the reliability of critical Azure-based healthcare services with an accent on SLOs, observability, automation, incident response, and cloud operations. Focus on designing resilient systems, implementing monitoring and alerting, improving disaster recovery, and driving measurable reliability improvements through post-incident analysis.
Location: Remote, any location in Argentina
Company
provides virtual care services and develops technology that expands access to healthcare.
What you will do
- Operate, scale, and improve the reliability of mission-critical Azure cloud services.
- Define and implement service level indicators, service level objectives, and error budget policies.
- Build observability across applications, infrastructure, networks, and cloud services using dashboards, alerts, logs, traces, and metrics.
- Analyze performance, capacity, resilience, backup, disaster recovery, failover, and business continuity.
- Participate in production readiness reviews and improve incident response, runbooks, and post-incident practices.
- Partner with engineering, product, security, cloud, network, operations, and incident management teams.
Requirements
- 7+ years of experience in site reliability engineering with ownership of mission-critical services.
- Deep Microsoft Azure experience, including Azure Monitor, Application Insights, AKS, and hybrid cloud-native operations.
- Production experience designing and rolling out SLI, SLO, and error budget programs.
- Hands-on experience with enterprise observability platforms such as Datadog, Dynatrace, Elastic, Grafana, Prometheus, or LogicMonitor.
- Hands-on experience implementing and configuring Datadog for monitoring, observability, and alerting.
Nice to have
- Experience establishing an SRE practice across multiple engineering teams.
- Strong incident command and blameless postmortem experience.
- Healthcare IT experience and familiarity with HIPAA, HITRUST, or equivalent frameworks.
- AWS reliability experience, chaos engineering, and resilience testing.
- Terraform, Bicep, Ansible, Python, PowerShell, or Go experience for automation and tooling.
Culture & Benefits
- High-performance, inclusive, and innovative work environment.
- Opportunities for career growth, leadership, and meaningful professional development.
- Benefits programs designed to support employees and their families.
- Culture that values diverse perspectives and continuous improvement.
Hiring process
- Identity and credential verification.
- Live or video interviews.
- Fraud and misrepresentation screening; falsified information results in disqualification.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
7 дней назад
Senior Site Reliability Engineer (Kubernetes)
7 дней назад
Senior Site Reliability Engineer (Azure)
11 часов назад
Site Reliability Engineer - Platform (Azure)
2 дня назад
Engineer III, Site Reliability (SRE)
6 часов назад
Senior DevOps Engineer, Applications (Azure)
131 250 - 175 000$
6 дней назад
Senior Site Reliability Engineer (Cloud Infrastructure)
150 000 - 172 000$