16 часов назад
Senior Staff Service Reliability and Operational Intelligence Engineer
187 945 - 269 503$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Staff Service Reliability and Operational Intelligence Engineer (Cloud Infrastructure, Observability, and AIOps): Defining reliability strategy and operating observability, SLO, incident-management, capacity, and disaster-recovery systems for cloud-managed SaaS products with on-premises customer components, with an accent on distributed systems, Kubernetes, and production resilience. Focus on building AIOps and self-healing workflows, leading high-severity incidents, designing failure experiments, and driving systemic reliability improvements across regions and services.
Location: Santa Clara, California, United States; travel up to 25%. Access to export-controlled technology may require U.S. Person status, a license, or a confirmed license exception.
Salary: $187,945–$269,503 USD base compensation, plus bonus, equity, and benefits.
Company
develops integrated quantum computing, networking, sensing, and security platforms delivered through major cloud providers and customer-site deployments.
What you will do
- Define the multi-year reliability and production-readiness strategy across development, pre-production, and production environments.
- Establish service ownership, New Service Introduction, observability, SLO, error-budget, and operational-readiness standards.
- Lead the architecture of shared observability platforms covering logs, metrics, traces, profiles, dashboards, alerting, synthetic monitoring, and telemetry governance.
- Command high-severity incidents, improve incident response and post-incident remediation, and drive systemic fixes for recurring failures.
- Lead capacity forecasting, performance testing, Kubernetes and cloud resource management, resilience exercises, disaster recovery, and validated failovers.
- Design AIOps and secure AI-agent workflows for event correlation, anomaly detection, autonomous triage, assisted remediation, and controlled self-healing.
Requirements
- 12+ years of production engineering, site reliability engineering, platform engineering, or cloud operations experience, including recent hands-on reliability work.
- Experience designing and operating large-scale, fault-tolerant systems on AWS or GCP, with deep knowledge of distributed systems, Kubernetes, networking, CI/CD, and production failure modes.
- Hands-on ownership of observability architecture, instrumentation, SLIs, SLOs, error budgets, and measurable reliability outcomes.
- Experience designing failure experiments, disaster-recovery exercises, service failovers, capacity strategies, and performance testing.
- Experience commanding SEV1 or SEV2 incidents and driving root causes through systemic remediation.
- Strong Python or Go software engineering, infrastructure-as-code, automation, architecture leadership, coaching, and cross-functional influence skills.
Nice to have
- Experience with AIOps, autonomous remediation, self-healing workflows, and governed AI agents integrated with operational platforms.
- Experience with Amazon Bedrock AgentCore or comparable agentic automation frameworks.
- Knowledge of FinOps, capacity optimization, telemetry cost management, global traffic management, progressive delivery, and follow-the-sun on-call models.
Culture & Benefits
- Autonomy-focused environment emphasizing productivity, respect, inclusion, and equal opportunity.
- Medical, dental, and vision coverage with matching 401(k).
- Unlimited paid time off, paid holidays, and parental/adoption leave.
- Legal insurance and a home technology stipend.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
10 часов назад
Senior DevOps Engineer / Site Reliability Engineer (AI)
170 000 - 220 000$
6 дней назад
Site Reliability Engineer (Kubernetes)
130 000 - 160 000$
SandboxAQ
6 дней назад
Staff Platform Engineer (AI)
121 600 - 228 000$
13 часов назад
Staff Cloud Reliability Engineer
180 000 - 200 000$
8 часов назад
Lead DevOps Engineer (AWS)
180 000 - 210 000$
15 часов назад
Staff Site Reliability Engineer (Infrastructure)
217 565 - 260 000$