Назад
Company hidden
6 дней назад

Staff Site Reliability Engineer (AI Ops)

152 000 - 228 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Staff Site Reliability Engineer (AI Ops): Building and operating reliable infrastructure for cloud-managed SaaS products with on-premises customer components, with an accent on observability, SLO governance, resilience, disaster recovery, and cloud security. Focus on leading high-severity incident response, designing chaos and failover automation, improving stateful and streaming platform reliability, and creating AI Ops self-healing workflows.

Location: Santa Clara, California, United States office, with the option to work remotely a few days per week. Travel up to 25%. Access to export-controlled technology may require U.S. Person status, an applicable license, or a license exception.

Salary: $152,000–$228,000 USD base compensation, plus bonus, equity, and benefits.

Company

hirify.global develops quantum computing, networking, sensing, and security platforms and delivers quantum services through major cloud providers.

What you will do

  • Set technical direction for production reliability across regions and services, including SLOs, error budgets, standards, and escalation policies.
  • Design and operate observability platforms covering metrics, logs, distributed tracing, and profiling.
  • Lead chaos engineering, disaster-recovery testing, failover validation, and resilience improvements against defined recovery objectives.
  • Command high-severity incidents, coordinate executive communication, and drive blameless post-incident systemic fixes.
  • Own reliability for Postgres, Redis/Valkey, Kafka, OpenSearch, and other stateful and streaming services, including capacity and efficiency engineering.
  • Build AI Ops workflows for predictive alerting, autonomous triage, remediation, and self-healing while mentoring engineers and aligning partner teams.

Requirements

  • 7+ years of production engineering experience with recent hands-on reliability work.
  • Recent experience operating large-scale, fault-tolerant production systems on AWS or GCP.
  • Experience instrumenting production systems and governing SLOs and error budgets.
  • Experience designing and executing failure experiments or disaster-recovery exercises with real failover validation.
  • Personal experience commanding serious SEV1/SEV2 incidents and driving root cause analysis through systemic remediation.
  • Evidence of measurable reliability outcomes and multi-team technical leadership through standards, reviews, coaching, and adopted mechanisms.

Nice to have

  • Cloud security posture management, runtime vulnerability detection, workload protection, and risk-based remediation experience.
  • AIOps, autonomous remediation, self-healing workflows, or Amazon Bedrock Agent Core experience.
  • Capacity management, resource rightsizing, FinOps-based cost optimization, and load-balancing operations.
  • LLM gateway traffic management, including routing, policy enforcement, rate limiting, model fallback, latency optimization, cost controls, and multi-model observability.
  • Ability to connect networking, security, and reliability concerns in platform design.

Culture & Benefits

  • Hybrid work with office-based collaboration in Santa Clara and limited remote work during the week.
  • Comprehensive medical, dental, and vision plans.
  • Matching 401(k), unlimited PTO, and paid holidays.
  • Parental and adoption leave, legal insurance, and a home technology stipend.
  • Autonomy, productivity, respect, and an inclusive work environment.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →