Назад
Company hidden
7 часов назад

Senior Site Reliability Engineer (Azure)

Формат работы
remote (только Argentina)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Argentina
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Site Reliability Engineer (Azure/Observability): Operating, scaling, and improving the reliability of critical Azure-based healthcare services with an accent on SLOs, observability, automation, incident response, and cloud operations. Focus on designing resilient systems, implementing monitoring and alerting, improving disaster recovery, and driving measurable reliability improvements through post-incident analysis.

Location: Remote, any location in Argentina

Company

hirify.global provides virtual care services and develops technology that expands access to healthcare.

What you will do

  • Operate, scale, and improve the reliability of mission-critical Azure cloud services.
  • Define and implement service level indicators, service level objectives, and error budget policies.
  • Build observability across applications, infrastructure, networks, and cloud services using dashboards, alerts, logs, traces, and metrics.
  • Analyze performance, capacity, resilience, backup, disaster recovery, failover, and business continuity.
  • Participate in production readiness reviews and improve incident response, runbooks, and post-incident practices.
  • Partner with engineering, product, security, cloud, network, operations, and incident management teams.

Requirements

  • 7+ years of experience in site reliability engineering with ownership of mission-critical services.
  • Deep Microsoft Azure experience, including Azure Monitor, Application Insights, AKS, and hybrid cloud-native operations.
  • Production experience designing and rolling out SLI, SLO, and error budget programs.
  • Hands-on experience with enterprise observability platforms such as Datadog, Dynatrace, Elastic, Grafana, Prometheus, or LogicMonitor.
  • Hands-on experience implementing and configuring Datadog for monitoring, observability, and alerting.

Nice to have

  • Experience establishing an SRE practice across multiple engineering teams.
  • Strong incident command and blameless postmortem experience.
  • Healthcare IT experience and familiarity with HIPAA, HITRUST, or equivalent frameworks.
  • AWS reliability experience, chaos engineering, and resilience testing.
  • Terraform, Bicep, Ansible, Python, PowerShell, or Go experience for automation and tooling.

Culture & Benefits

  • High-performance, inclusive, and innovative work environment.
  • Opportunities for career growth, leadership, and meaningful professional development.
  • Benefits programs designed to support employees and their families.
  • Culture that values diverse perspectives and continuous improvement.

Hiring process

  • Identity and credential verification.
  • Live or video interviews.
  • Fraud and misrepresentation screening; falsified information results in disqualification.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →