3 часа назад
Senior DevOps Engineer / Site Reliability Engineer (AI)
170 000 - 220 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior DevOps Engineer / Site Reliability Engineer (AI) (Kubernetes, Terraform, GCP, Azure): Ensuring the reliability and operational readiness of production infrastructure powering AI solutions for major health systems with an accent on observability, incident response, secure multi-cloud networking, and deployment automation. Focus on managing Kubernetes and Temporal workloads, designing CI/CD pipelines, codifying operations with Python and Terraform, and maintaining HIPAA and HITRUST controls.
Location: Palo Alto, United States — hybrid
Salary: $170,000–$220,000 per year, plus equity and benefits
Company
builds healthcare infrastructure for safe generative AI governance, healthcare-specific agents, and real-time algorithm monitoring.
What you will do
- Partner with engineering teams to assess production readiness, deployment patterns, failure modes, resource requirements, and rollback strategies.
- Design and maintain observability infrastructure covering metrics, logging, distributed tracing, dashboards, and alerting across multi-cloud environments.
- Define SLIs/SLOs, manage on-call rotations, lead incident response, coordinate hotfixes, and drive root cause analysis.
- Build release documentation, runbooks, postmortems, operational playbooks, and developer-experience improvements.
- Operate Kubernetes and Temporal workloads in production, including troubleshooting, scaling, and resource optimization.
- Design secure zero-trust network architectures and improve CI/CD pipelines using Python and Terraform.
Requirements
- 6+ years of experience in SRE, DevOps, or infrastructure engineering, including at least 3 years managing production workloads.
- Strong Terraform and Python skills, with experience in module development, state management, automation, and operational tooling.
- Deep production Kubernetes experience and hands-on experience with Google Cloud Platform and Microsoft Azure.
- Strong networking and security knowledge, including zero-trust architecture, segmentation, private connectivity, identity-based access controls, and secrets management.
- Production experience with Temporal or comparable workflow orchestration systems, plus observability and incident-response ownership.
- Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent experience.
Nice to have
- Healthcare experience and familiarity with HIPAA, HITRUST, or similar compliance frameworks.
- Experience operating LLM-based systems, agentic workflows, or RAG pipelines in production.
- Experience with GitOps, multi-tenant SaaS infrastructure, chaos engineering, or reliability testing.
- Prior experience as a founding or early SRE/platform hire at a startup.
Culture & Benefits
- Mission-driven work focused on improving healthcare with AI.
- Competitive salary, equity package, and medical, dental, and vision insurance.
- Flexible working hours and hybrid work options.
- Inclusive environment focused on creativity, innovation, and diverse perspectives.
Hiring process
- Submit a resume through the application process.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
5 часов назад
Site Reliability Engineer (AI)
100 000 - 300 000$
5 часов назад
Staff Site Reliability Engineer (AI/ML)
241 000 - 270 000$
10 часов назад
DevOps Engineer (AI)
85 000 - 180 000$
6 часов назад
Infrastructure Engineer (AI)
160 000 - 245 000$
7 часов назад
Senior Platform Engineer (AI)
170 000 - 240 000$
6 часов назад
Staff Software Engineer (AWS/Kubernetes)
175 000 - 245 000$