Назад
Company hidden
1 день назад

Engineer III, Site Reliability (SRE)

Формат работы
remote (только USA)/hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Engineer III, Site Reliability (SRE) (Cloud Operations/Kubernetes): Building and operating reliable, scalable cloud services and delivery platforms for mission-critical pharmacy automation systems with an accent on observability, automation, and incident response. Focus on defining SLIs and SLOs, designing Terraform-based infrastructure and CI/CD pipelines, and implementing ML-based anomaly detection and automated diagnostics.

Location: Hybrid or remote within the United States; up to 10% travel and participation in an SRE on-call rotation required.

Company

hirify.global develops cloud-native SaaS solutions for managing medications and supplies across the healthcare continuum.

What you will do

  • Own the reliability, scalability, instrumentation, alerting, dashboards, and runbooks for assigned cloud services.
  • Define SLIs and SLOs with product and engineering teams and drive improvements in observability, automation, and resilience.
  • Participate in on-call operations, lead Sev-2 and Sev-3 incident response as readiness grows, and support Sev-1 technical leadership.
  • Lead blameless post-incident reviews and track follow-up actions to completion.
  • Design and operate CI/CD pipelines using GitHub Actions, CodeFresh, TeamCity, or Octopus Deploy.
  • Automate infrastructure with Terraform and contribute to observability, golden paths, architecture reviews, and launch readiness.

Requirements

  • Bachelor’s degree in Computer Science, Engineering, or a related technical field.
  • 5+ years in software or platform engineering, including 3+ years in SRE, DevOps, or reliability-focused roles.
  • Hands-on experience with AWS, Azure, or GCP, plus Python or another object-oriented programming language.
  • Production experience with Kubernetes, Docker, Helm, and Terraform or similar Infrastructure as Code frameworks.
  • Knowledge of metrics, logs, tracing, Linux administration, incident response, on-call operations, and post-incident write-ups.
  • Collaborative and coachable approach with an interest in developing under senior SRE mentorship.

Nice to have

  • Experience in regulated environments such as healthcare, financial services, or government, including HIPAA or SOC 2.
  • Exposure to managed service provider models, AIOps, ML-based anomaly detection, or LLM-assisted incident triage.
  • Knowledge of GitOps tools such as ArgoCD or Flux and secure, compliant Kubernetes platforms.
  • Experience with chaos engineering, Kafka, RabbitMQ, or stateful Kubernetes services.

Culture & Benefits

  • Remote and hybrid work environments are supported within the United States.
  • Work in a newly formed SRE practice with direct mentorship from a Senior SRE.
  • Collaborate with product engineering, security, operations, and managed service providers.
  • Opportunities for progression to Senior Site Reliability Engineer and lateral growth into platform, security, or product engineering.
  • Up to 10% travel as needed.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →