3 дня назад
Sr. Manager, Site Reliability
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Sr. Manager, Site Reliability (SRE, Kubernetes, AIOps): Building Omnicell's cloud reliability practice for hybrid hardware-plus-cloud and SaaS healthcare platforms with an accent on SLOs, incident command, observability, and compliant operations. Focus on designing the SRE operating model, leading hands-on incident response and Kubernetes infrastructure work, and introducing explainable AI-assisted monitoring in HIPAA and SOC 2 environments.
Location: Austin, Texas, United States; hybrid environment with remote work supported. Travel up to 10% and on-call participation are expected.
Company
develops medication-dispensing products and cloud services used by hospitals, including hardware-connected and SaaS platforms.
What you will do
- Build the SRE practice from zero, defining SLOs, SLIs, error-budget policies, incident severity standards, and postmortem processes.
- Select and establish the observability platform, instrumentation standards, operational KPIs, and executive reliability reporting.
- Instrument Tier-1 services, create dashboards, alerts, and runbooks, participate in on-call, and command Sev-1 and Sev-2 incidents.
- Contribute Python automation, Terraform infrastructure-as-code, CI/CD pipelines, Kubernetes platform improvements, and chaos or failover exercises.
- Architect and evaluate AIOps capabilities including anomaly detection, predictive alerting, automated root-cause analysis, and LLM-assisted runbooks.
- Coach an Engineer III SRE, design future SRE hiring plans, and partner with Engineering, Security, Compliance, Architecture, Product, Support, IBM, and HCL.
Requirements
- 8+ years in software or platform engineering, including at least 4 years in SRE, DevOps, or platform reliability.
- At least 2 years of technical leadership, tech-lead, or staff-level experience with mentorship responsibilities.
- Experience establishing SLOs, incident command, error budgets, observability practices, and reliability operations from zero or near-zero.
- Deep hands-on experience with AWS, Azure, or GCP, plus networking, IAM, managed services, Kubernetes, Docker, Helm, and service mesh technologies.
- Strong experience with CI/CD, Terraform or comparable infrastructure-as-code tools, Python or another object-oriented language, and modern observability platforms.
- Bachelor's degree in Computer Science, Engineering, or a related technical field, or equivalent experience; experience in regulated environments such as HIPAA or SOC 2.
Nice to have
- Healthcare or other regulated-industry experience and experience with hybrid hardware-plus-cloud products.
- Experience with AIOps, LLM APIs, agentic AI frameworks, stateful Kubernetes services, chaos engineering, Kafka, RabbitMQ, or security monitoring.
- Experience integrating managed service providers into an internal SRE operating model and applying FinOps or cloud cost optimization practices.
Culture & Benefits
- Hybrid work in a corporate office or lab environment, with remote work supported.
- Hands-on player-coach role with approximately 60% engineering and incident response, 25% practice design and coaching, and 15% cross-functional partnership.
- Opportunity to shape reliability standards and AI-driven operations across a healthcare platform where service availability affects pharmacy workflows and patient care.
- On-call participation is part of the SRE rotation, with a compensation approach and fair distribution to be established.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
22 часа назад
Manager of Site Reliability Engineering (Azure/AWS)
160 000 - 180 000$
2 дня назад
Senior Manager, Engineering (DevOps, Infrastructure, and Release Engineering)
5 дней назад
Engineering Manager - DevOps
125 000 - 175 000$
12 часов назад
Sr. Manager Database Platform (SQL Server)
136 885 - 200 000$
4 дня назад
Sr. Manager, Site Reliability Engineering (DevOps)
150 000 - 170 000$
8 часов назад
Engineering Manager (AI)
210 000 - 300 000$