Назад
Company hidden
7 дней назад

Staff Service Reliability and Operational Intelligence Engineer (AI Ops)

Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Staff Service Reliability and Operational Intelligence Engineer (AI Ops): Building and operating resilient observability, reliability, and AI Ops capabilities for cloud-managed SaaS products with on-premises customer deployments, with an accent on SLO governance, incident response, capacity management, and disaster recovery. Focus on designing self-healing workflows, leading high-severity incidents, and solving complex distributed-systems and production-resilience challenges.

Location: Santa Clara, California, United States; hybrid schedule with a few remote days per week. Travel up to 25%. Access to export-controlled technology requires U.S. Person status, an applicable license, or a confirmed license exception.

Company

hirify.global develops quantum computing platforms and integrated quantum solutions for computing, networking, sensing, and security.

What you will do

  • Define the technical strategy and multi-year roadmap for operational excellence, production readiness, and service reliability.
  • Build and govern shared observability platforms covering logs, metrics, distributed traces, profiles, dashboards, alerts, synthetic monitoring, and telemetry quality.
  • Establish service ownership, SLI/SLO, error-budget, incident-management, escalation, and on-call standards across production services.
  • Lead high-severity incident response, blameless post-incident reviews, systemic remediation, resilience exercises, disaster recovery, and validated failovers.
  • Manage capacity forecasting, performance testing, Kubernetes and cloud resources, scaling thresholds, rightsizing, and operational efficiency.
  • Design secure AI Ops and AI-agent workflows for event correlation, predictive detection, root-cause analysis, triage, remediation, and controlled self-healing.

Requirements

  • 8+ years of production engineering, site reliability engineering, platform engineering, or cloud operations experience, including recent hands-on reliability work.
  • Experience designing and operating large-scale fault-tolerant systems on AWS or GCP.
  • Deep knowledge of distributed systems, cloud infrastructure, Kubernetes, networking, CI/CD, and production failure modes.
  • Hands-on experience with observability architecture, instrumentation, metrics, logs, traces, SLIs, SLOs, and error budgets.
  • Experience commanding SEV1 or SEV2 incidents, conducting failure experiments and disaster-recovery exercises, and driving systemic remediation.
  • Strong automation and software engineering skills with Python or Go, infrastructure as code, and modern delivery toolchains; ability to lead across teams without direct authority.

Nice to have

  • Experience designing AI Ops, autonomous remediation, self-healing workflows, and governed AI-agent integrations with operational platforms.
  • Experience with Amazon Bedrock AgentCore or comparable agentic automation frameworks.
  • Knowledge of FinOps, telemetry cost management, capacity optimization, load balancing, global traffic management, and highly available service design.
  • Experience with canary deployments, blue-green deployments, automated rollback, feature flags, and global follow-the-sun on-call models.

Culture & Benefits

  • Autonomy, productivity, respect, and an inclusive environment focused on removing barriers.
  • Medical, dental, and vision coverage with matching 401(k).
  • Unlimited paid time off, paid holidays, and parental/adoption leave.
  • Legal insurance and a home technology stipend.
  • Total compensation includes base salary, bonus, equity, and benefits.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →