6 дней назад
Staff Site Reliability Engineer (AI Ops)
152 000 - 228 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Staff Site Reliability Engineer (AI Ops): Building and operating reliable infrastructure for cloud-managed SaaS products with on-premises customer components, with an accent on observability, SLO governance, resilience, disaster recovery, and cloud security. Focus on leading high-severity incident response, designing chaos and failover automation, improving stateful and streaming platform reliability, and creating AI Ops self-healing workflows.
Location: Santa Clara, California, United States office, with the option to work remotely a few days per week. Travel up to 25%. Access to export-controlled technology may require U.S. Person status, an applicable license, or a license exception.
Salary: $152,000–$228,000 USD base compensation, plus bonus, equity, and benefits.
Company
develops quantum computing, networking, sensing, and security platforms and delivers quantum services through major cloud providers.
What you will do
- Set technical direction for production reliability across regions and services, including SLOs, error budgets, standards, and escalation policies.
- Design and operate observability platforms covering metrics, logs, distributed tracing, and profiling.
- Lead chaos engineering, disaster-recovery testing, failover validation, and resilience improvements against defined recovery objectives.
- Command high-severity incidents, coordinate executive communication, and drive blameless post-incident systemic fixes.
- Own reliability for Postgres, Redis/Valkey, Kafka, OpenSearch, and other stateful and streaming services, including capacity and efficiency engineering.
- Build AI Ops workflows for predictive alerting, autonomous triage, remediation, and self-healing while mentoring engineers and aligning partner teams.
Requirements
- 7+ years of production engineering experience with recent hands-on reliability work.
- Recent experience operating large-scale, fault-tolerant production systems on AWS or GCP.
- Experience instrumenting production systems and governing SLOs and error budgets.
- Experience designing and executing failure experiments or disaster-recovery exercises with real failover validation.
- Personal experience commanding serious SEV1/SEV2 incidents and driving root cause analysis through systemic remediation.
- Evidence of measurable reliability outcomes and multi-team technical leadership through standards, reviews, coaching, and adopted mechanisms.
Nice to have
- Cloud security posture management, runtime vulnerability detection, workload protection, and risk-based remediation experience.
- AIOps, autonomous remediation, self-healing workflows, or Amazon Bedrock Agent Core experience.
- Capacity management, resource rightsizing, FinOps-based cost optimization, and load-balancing operations.
- LLM gateway traffic management, including routing, policy enforcement, rate limiting, model fallback, latency optimization, cost controls, and multi-model observability.
- Ability to connect networking, security, and reliability concerns in platform design.
Culture & Benefits
- Hybrid work with office-based collaboration in Santa Clara and limited remote work during the week.
- Comprehensive medical, dental, and vision plans.
- Matching 401(k), unlimited PTO, and paid holidays.
- Parental and adoption leave, legal insurance, and a home technology stipend.
- Autonomy, productivity, respect, and an inclusive work environment.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
13 дней назад
Staff Site Reliability Engineer (AI)
252 000 - 308 000$
8 дней назад
Sr. Site Reliability Engineer
160 000 - 180 000$
7 дней назад
Staff Software Engineer, Reliability
203 500 - 248 500$
Vapi
12 дней назад
Senior Site Reliability Engineer (AI)
280 000 - 314 000$
7 дней назад
Senior Site Reliability Engineer (AWS)
12 дней назад
Senior Site Reliability Engineer (AI)
152 500 - 205 000$