1 день назад
Engineer III, Site Reliability (SRE)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Engineer III, Site Reliability (SRE) (Cloud Operations/Kubernetes): Building and operating reliable, scalable cloud services and delivery platforms for mission-critical pharmacy automation systems with an accent on observability, automation, and incident response. Focus on defining SLIs and SLOs, designing Terraform-based infrastructure and CI/CD pipelines, and implementing ML-based anomaly detection and automated diagnostics.
Location: Hybrid or remote within the United States; up to 10% travel and participation in an SRE on-call rotation required.
Company
develops cloud-native SaaS solutions for managing medications and supplies across the healthcare continuum.
What you will do
- Own the reliability, scalability, instrumentation, alerting, dashboards, and runbooks for assigned cloud services.
- Define SLIs and SLOs with product and engineering teams and drive improvements in observability, automation, and resilience.
- Participate in on-call operations, lead Sev-2 and Sev-3 incident response as readiness grows, and support Sev-1 technical leadership.
- Lead blameless post-incident reviews and track follow-up actions to completion.
- Design and operate CI/CD pipelines using GitHub Actions, CodeFresh, TeamCity, or Octopus Deploy.
- Automate infrastructure with Terraform and contribute to observability, golden paths, architecture reviews, and launch readiness.
Requirements
- Bachelor’s degree in Computer Science, Engineering, or a related technical field.
- 5+ years in software or platform engineering, including 3+ years in SRE, DevOps, or reliability-focused roles.
- Hands-on experience with AWS, Azure, or GCP, plus Python or another object-oriented programming language.
- Production experience with Kubernetes, Docker, Helm, and Terraform or similar Infrastructure as Code frameworks.
- Knowledge of metrics, logs, tracing, Linux administration, incident response, on-call operations, and post-incident write-ups.
- Collaborative and coachable approach with an interest in developing under senior SRE mentorship.
Nice to have
- Experience in regulated environments such as healthcare, financial services, or government, including HIPAA or SOC 2.
- Exposure to managed service provider models, AIOps, ML-based anomaly detection, or LLM-assisted incident triage.
- Knowledge of GitOps tools such as ArgoCD or Flux and secure, compliant Kubernetes platforms.
- Experience with chaos engineering, Kafka, RabbitMQ, or stateful Kubernetes services.
Culture & Benefits
- Remote and hybrid work environments are supported within the United States.
- Work in a newly formed SRE practice with direct mentorship from a Senior SRE.
- Collaborate with product engineering, security, operations, and managed service providers.
- Opportunities for progression to Senior Site Reliability Engineer and lateral growth into platform, security, or product engineering.
- Up to 10% travel as needed.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
3 дня назад
Site Reliability Engineer (Cybersecurity)
96 500 - 183 500$
22 часа назад
Senior Site Reliability Engineer (AWS)
6 дней назад
Senior Site Reliability Engineer (Fintech)
160 000 - 200 000$
2 часа назад
Senior SecDevOps Engineer (AI)
165 000 - 200 000$
50 минут назад
Sr Site Reliability Engineer (AWS)
95 000 - 135 000$
3 дня назад