Назад
Company hidden
16 часов назад

Senior Staff Service Reliability and Operational Intelligence Engineer

187 945 - 269 503$
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
UK/US/SK +4 еще
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Staff Service Reliability and Operational Intelligence Engineer (Cloud Infrastructure, Observability, and AIOps): Defining reliability strategy and operating observability, SLO, incident-management, capacity, and disaster-recovery systems for cloud-managed SaaS products with on-premises customer components, with an accent on distributed systems, Kubernetes, and production resilience. Focus on building AIOps and self-healing workflows, leading high-severity incidents, designing failure experiments, and driving systemic reliability improvements across regions and services.

Location: Santa Clara, California, United States; travel up to 25%. Access to export-controlled technology may require U.S. Person status, a license, or a confirmed license exception.

Salary: $187,945–$269,503 USD base compensation, plus bonus, equity, and benefits.

Company

hirify.global develops integrated quantum computing, networking, sensing, and security platforms delivered through major cloud providers and customer-site deployments.

What you will do

  • Define the multi-year reliability and production-readiness strategy across development, pre-production, and production environments.
  • Establish service ownership, New Service Introduction, observability, SLO, error-budget, and operational-readiness standards.
  • Lead the architecture of shared observability platforms covering logs, metrics, traces, profiles, dashboards, alerting, synthetic monitoring, and telemetry governance.
  • Command high-severity incidents, improve incident response and post-incident remediation, and drive systemic fixes for recurring failures.
  • Lead capacity forecasting, performance testing, Kubernetes and cloud resource management, resilience exercises, disaster recovery, and validated failovers.
  • Design AIOps and secure AI-agent workflows for event correlation, anomaly detection, autonomous triage, assisted remediation, and controlled self-healing.

Requirements

  • 12+ years of production engineering, site reliability engineering, platform engineering, or cloud operations experience, including recent hands-on reliability work.
  • Experience designing and operating large-scale, fault-tolerant systems on AWS or GCP, with deep knowledge of distributed systems, Kubernetes, networking, CI/CD, and production failure modes.
  • Hands-on ownership of observability architecture, instrumentation, SLIs, SLOs, error budgets, and measurable reliability outcomes.
  • Experience designing failure experiments, disaster-recovery exercises, service failovers, capacity strategies, and performance testing.
  • Experience commanding SEV1 or SEV2 incidents and driving root causes through systemic remediation.
  • Strong Python or Go software engineering, infrastructure-as-code, automation, architecture leadership, coaching, and cross-functional influence skills.

Nice to have

  • Experience with AIOps, autonomous remediation, self-healing workflows, and governed AI agents integrated with operational platforms.
  • Experience with Amazon Bedrock AgentCore or comparable agentic automation frameworks.
  • Knowledge of FinOps, capacity optimization, telemetry cost management, global traffic management, progressive delivery, and follow-the-sun on-call models.

Culture & Benefits

  • Autonomy-focused environment emphasizing productivity, respect, inclusion, and equal opportunity.
  • Medical, dental, and vision coverage with matching 401(k).
  • Unlimited paid time off, paid holidays, and parental/adoption leave.
  • Legal insurance and a home technology stipend.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →