7 дней назад
Staff Service Reliability and Operational Intelligence Engineer (AI Ops)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Staff Service Reliability and Operational Intelligence Engineer (AI Ops): Building and operating resilient observability, reliability, and AI Ops capabilities for cloud-managed SaaS products with on-premises customer deployments, with an accent on SLO governance, incident response, capacity management, and disaster recovery. Focus on designing self-healing workflows, leading high-severity incidents, and solving complex distributed-systems and production-resilience challenges.
Location: Santa Clara, California, United States; hybrid schedule with a few remote days per week. Travel up to 25%. Access to export-controlled technology requires U.S. Person status, an applicable license, or a confirmed license exception.
Company
develops quantum computing platforms and integrated quantum solutions for computing, networking, sensing, and security.
What you will do
- Define the technical strategy and multi-year roadmap for operational excellence, production readiness, and service reliability.
- Build and govern shared observability platforms covering logs, metrics, distributed traces, profiles, dashboards, alerts, synthetic monitoring, and telemetry quality.
- Establish service ownership, SLI/SLO, error-budget, incident-management, escalation, and on-call standards across production services.
- Lead high-severity incident response, blameless post-incident reviews, systemic remediation, resilience exercises, disaster recovery, and validated failovers.
- Manage capacity forecasting, performance testing, Kubernetes and cloud resources, scaling thresholds, rightsizing, and operational efficiency.
- Design secure AI Ops and AI-agent workflows for event correlation, predictive detection, root-cause analysis, triage, remediation, and controlled self-healing.
Requirements
- 8+ years of production engineering, site reliability engineering, platform engineering, or cloud operations experience, including recent hands-on reliability work.
- Experience designing and operating large-scale fault-tolerant systems on AWS or GCP.
- Deep knowledge of distributed systems, cloud infrastructure, Kubernetes, networking, CI/CD, and production failure modes.
- Hands-on experience with observability architecture, instrumentation, metrics, logs, traces, SLIs, SLOs, and error budgets.
- Experience commanding SEV1 or SEV2 incidents, conducting failure experiments and disaster-recovery exercises, and driving systemic remediation.
- Strong automation and software engineering skills with Python or Go, infrastructure as code, and modern delivery toolchains; ability to lead across teams without direct authority.
Nice to have
- Experience designing AI Ops, autonomous remediation, self-healing workflows, and governed AI-agent integrations with operational platforms.
- Experience with Amazon Bedrock AgentCore or comparable agentic automation frameworks.
- Knowledge of FinOps, telemetry cost management, capacity optimization, load balancing, global traffic management, and highly available service design.
- Experience with canary deployments, blue-green deployments, automated rollback, feature flags, and global follow-the-sun on-call models.
Culture & Benefits
- Autonomy, productivity, respect, and an inclusive environment focused on removing barriers.
- Medical, dental, and vision coverage with matching 401(k).
- Unlimited paid time off, paid holidays, and parental/adoption leave.
- Legal insurance and a home technology stipend.
- Total compensation includes base salary, bonus, equity, and benefits.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
10 дней назад
Site Reliability Engineer (Cloud Banking)
9 дней назад
Site Reliability Engineer, IaaS (AI)
8 дней назад
Principal Production Engineer (AI)
164 500 - 235 000$
9 дней назад
Sr. Site Reliability Engineer
160 000 - 180 000$
7 дней назад
Site Reliability Engineering Team Lead (Principal SRE, Automotive AI)
13 дней назад
Senior Site Reliability Engineer (AI)
152 500 - 205 000$