Назад
Company hidden
3 дня назад

Sr. Manager, Site Reliability

Формат работы
remote (только USA)/hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Sr. Manager, Site Reliability (SRE, Kubernetes, AIOps): Building Omnicell's cloud reliability practice for hybrid hardware-plus-cloud and SaaS healthcare platforms with an accent on SLOs, incident command, observability, and compliant operations. Focus on designing the SRE operating model, leading hands-on incident response and Kubernetes infrastructure work, and introducing explainable AI-assisted monitoring in HIPAA and SOC 2 environments.

Location: Austin, Texas, United States; hybrid environment with remote work supported. Travel up to 10% and on-call participation are expected.

Company

hirify.global develops medication-dispensing products and cloud services used by hospitals, including hardware-connected and SaaS platforms.

What you will do

  • Build the SRE practice from zero, defining SLOs, SLIs, error-budget policies, incident severity standards, and postmortem processes.
  • Select and establish the observability platform, instrumentation standards, operational KPIs, and executive reliability reporting.
  • Instrument Tier-1 services, create dashboards, alerts, and runbooks, participate in on-call, and command Sev-1 and Sev-2 incidents.
  • Contribute Python automation, Terraform infrastructure-as-code, CI/CD pipelines, Kubernetes platform improvements, and chaos or failover exercises.
  • Architect and evaluate AIOps capabilities including anomaly detection, predictive alerting, automated root-cause analysis, and LLM-assisted runbooks.
  • Coach an Engineer III SRE, design future SRE hiring plans, and partner with Engineering, Security, Compliance, Architecture, Product, Support, IBM, and HCL.

Requirements

  • 8+ years in software or platform engineering, including at least 4 years in SRE, DevOps, or platform reliability.
  • At least 2 years of technical leadership, tech-lead, or staff-level experience with mentorship responsibilities.
  • Experience establishing SLOs, incident command, error budgets, observability practices, and reliability operations from zero or near-zero.
  • Deep hands-on experience with AWS, Azure, or GCP, plus networking, IAM, managed services, Kubernetes, Docker, Helm, and service mesh technologies.
  • Strong experience with CI/CD, Terraform or comparable infrastructure-as-code tools, Python or another object-oriented language, and modern observability platforms.
  • Bachelor's degree in Computer Science, Engineering, or a related technical field, or equivalent experience; experience in regulated environments such as HIPAA or SOC 2.

Nice to have

  • Healthcare or other regulated-industry experience and experience with hybrid hardware-plus-cloud products.
  • Experience with AIOps, LLM APIs, agentic AI frameworks, stateful Kubernetes services, chaos engineering, Kafka, RabbitMQ, or security monitoring.
  • Experience integrating managed service providers into an internal SRE operating model and applying FinOps or cloud cost optimization practices.

Culture & Benefits

  • Hybrid work in a corporate office or lab environment, with remote work supported.
  • Hands-on player-coach role with approximately 60% engineering and incident response, 25% practice design and coaching, and 15% cross-functional partnership.
  • Opportunity to shape reliability standards and AI-driven operations across a healthcare platform where service availability affects pharmacy workflows and patient care.
  • On-call participation is part of the SRE rotation, with a compensation approach and fair distribution to be established.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →