3 часа назад
Senior Site Reliability Engineer (Systems Engineer III) (GCP)
140 000 - 150 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Site Reliability Engineer (Systems Engineer III) (GCP): Building and operating reliable, observable, and high-performing GCP infrastructure for a telehealth platform with an accent on SLOs, incident response, infrastructure as code, and HIPAA/SOC 2 controls. Focus on reducing latency and Cloud Run cold starts, expanding Terraform and CI/CD, testing disaster recovery, and using AI-assisted operational tooling.
Location: United States; remote work
Salary: $140,000–$150,000 USD annually
Company
is a venture-backed healthcare startup providing FDA-approved medication and clinical care for opioid use disorder through a telehealth platform.
What you will do
- Own reliability, availability, performance, SLOs, SLIs, and disaster recovery for the GCP platform.
- Build observability with Cloud Operations, log-based metrics, dashboards, alerting, and PagerDuty incident response.
- Expand Terraform-based infrastructure as code across GCP and improve the GitHub Actions CI/CD pipeline.
- Reduce latency, Cloud Run cold starts, and recovery time while testing Firestore backup, restore, and multi-region resilience.
- Partner with IT on IAM, secrets management, audit logging, and HIPAA/SOC 2 infrastructure controls.
- Enable fullstack engineers through documentation, runbooks, pairing, code review, and reliability guidance.
Requirements
- 5+ years of engineering experience, including 3+ years in cloud infrastructure, DevOps, or SRE.
- Expert hands-on GCP experience with Cloud Run, Cloud Operations, IAM, Firebase, and Firestore; Terraform fluency is strongly preferred.
- Expertise in at least one infrastructure automation language such as Python and strong shell scripting skills.
- Ability to read and debug TypeScript and Node application code; SQL for BigQuery and Log Analytics is a plus.
- Production operations experience with on-call rotations, incident response, postmortems, SLOs, error budgets, and quantitative monitoring.
- Must be based in the United States. Daily hands-on use of AI tools in engineering and operational automation is required.
Nice to have
- Healthcare, telehealth, or another regulated-industry background.
- Hands-on experience implementing HIPAA or SOC 2 infrastructure controls.
- Passion for expanding access to evidence-based addiction treatment.
Culture & Benefits
- Remote, growth-stage environment with autonomous ownership and proactive communication.
- Medical, vision, and health insurance, with many plans fully covered for employees.
- 20 days of PTO initially, increasing with tenure, plus 10 company holidays.
- Work-from-home stipend and 401(k) contribution platform.
- Additional life insurance, disability, financial wellness, and virtual primary care benefits.
Hiring process
- Applications are reviewed by a real person, with an emphasis on authentic candidate responses.
- Compensation follows a first-and-best offer approach, with the compensation bands discussed during interviews.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
2 дня назад
Sr. Manager, Site Reliability
59 550 - 110 594GBP
Okta
5 дней назад
Staff Site Reliability Engineer, Federal (TS/SCI)
174 000 - 238 000$
6 дней назад
Site Reliability Engineer (AI Infrastructure)
175 000 - 265 000$
Okta
5 дней назад
Staff Site Reliability Engineer (Splunk)
194 000 - 267 000$
Okta
5 дней назад
Staff Site Reliability Engineer (Kubernetes)
194 000 - 267 000$
6 дней назад
Site Reliability Engineer III (DBA)
125 000 - 150 000$