2 дня назад
Principal Site Reliability Engineer (AI)
165 000 - 185 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Principal Site Reliability Engineer (AI/SRE): Building reliable, observable, and compliant production platforms for diabetes technology products with an accent on incident management, infrastructure automation, disaster recovery, and AI-enabled tooling. Focus on designing SLOs, on-call systems, Terraform infrastructure, CI/CD guardrails, and cloud security while mentoring distributed engineering teams.
Location: Fully remote within the United States. Training is virtual and equipment is provided.
Salary: $165,000–$185,000 annually, plus bonus and benefits.
Company
develops and manufactures diabetes technology, including insulin pumps and automated insulin-delivery systems.
What you will do
- Lead production support, high-severity incident response, stakeholder communication, and blameless postmortems.
- Define SLIs and SLOs, improve observability, reduce MTTD and MTTR, and build runbooks and automation.
- Own on-call strategy, alert tuning, escalation paths, and support practices across distributed teams and time zones.
- Lead Terraform-based infrastructure automation, CI/CD reliability guardrails, technology lifecycle management, and toil reduction.
- Design and maintain disaster recovery, backup, failover, recovery testing, and RTO/RPO capabilities.
- Mentor SRE and DevOps engineers and partner with software engineering, QA, architecture, Security, Quality, and Compliance.
Requirements
- Demonstrated experience leading production support, incident management, and incident command.
- Strong knowledge of SRE practices, including SLOs, blameless postmortems, toil reduction, and on-call design.
- Expertise with Terraform or comparable infrastructure as code, including modules, state management, and policy guardrails.
- Hands-on experience with CI/CD, a major cloud platform, Docker, Kubernetes, observability tooling, and disaster recovery testing.
- Working knowledge of cloud security, compliance, IAM, vulnerability management, patching, and cloud cost optimization.
- Proficiency in Python, Go, Bash, or another scripting or programming language; experience using AI to improve internal tooling and integrations.
Nice to have
- 10+ years in SRE, DevOps, or infrastructure engineering and 2+ years mentoring or technically leading engineers.
- Experience in FDA- and ISO-regulated industries and with agile methodologies.
- Relevant cloud certifications and a computer science degree or equivalent practical experience.
Culture & Benefits
- People-first, human-centered work focused on improving diabetes care.
- Medical, dental, and vision coverage available from the first day.
- Health savings and flexible spending accounts, 401(k) matching, and an employee stock purchase plan.
- 11 paid holidays and at least 20 days of paid time off with accrual starting on day one.
- Inclusive workplace with virtual training and distributed collaboration.
Hiring process
- Employment is contingent on successful drug testing and background screening.
- Criminal history review may be conducted because the role can access sensitive and protected health information.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
5 дней назад
Senior DevOps / Site Reliability Engineer (SRE) (Cybersecurity)
165 000 - 215 000$
4 дня назад
Principal Site Reliability Engineer
220 000 - 280 000PLN
4 дня назад
Sr. Site Reliability Engineer
160 000 - 180 000$
2 дня назад
Site Reliability Engineering Team Lead (Principal SRE, Automotive AI)
Okta
3 дня назад
Staff TDI Site Reliability Engineer (AWS)
174 000 - 239 000$
Okta
3 дня назад
Staff TDI Site Reliability Engineer (Okta Federal)
174 000 - 239 000$