Назад
Company hidden
3 дня назад

Staff Site Reliability Engineer (AI)

252 000 - 308 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Staff Site Reliability Engineer (AI) (AWS/Kubernetes): Defining reliability standards and building AI-assisted operational workflows for critical financial services with an accent on incident response, observability, resilience, and production readiness. Focus on designing AI-driven alert triage and runbook automation, leading high-severity incidents, reducing on-call toil, and guiding large-scale AWS architecture across EKS, Kafka, DynamoDB, RDS, and SQS.

Location: Mountain View, US; hybrid position requiring in-office work 2 days a week

Salary: $252,000–$308,000 base salary annually, plus equity and benefits

Company

hirify.global builds earned wage access products that provide real-time financial flexibility, enabling members to access, spend, save, and grow their hirify.globalgs.

What you will do

  • Define reliability standards across critical services, including SLIs, SLOs, error budgets, observability, production readiness, incident response, and resilience patterns.
  • Build AI-assisted workflows for alert correlation, incident triage, root-cause exploration, runbook retrieval, postmortem drafting, and corrective-action tracking.
  • Lead high-severity incidents as Incident Commander and improve detection, response quality, communication, and recurrence prevention.
  • Develop tools that reduce on-call toil by gathering context from Datadog, CloudWatch, incident.io, Slack, runbooks, deployment history, and service metadata.
  • Guide resilient service architecture across AWS, including graceful degradation, failure isolation, capacity planning, and operational safety for EKS, Kafka, DynamoDB, RDS, and SQS.
  • Mentor engineers and influence product engineering, infrastructure, security, and leadership teams through design reviews, incident reviews, documentation, and reusable reliability practices.

Requirements

  • 7+ years of experience in SRE, software engineering, or infrastructure engineering with increasing scope and cross-organizational influence.
  • Demonstrated success improving reliability and operational excellence at scale using KPIs such as MTTR, MTTD, alert quality, incident recurrence, SLO attainment, on-call health, and corrective-action completion.
  • Experience applying AI or LLMs to engineering and operational workflows, plus practical use of AI-assisted development tools such as Cursor, Claude Code, Copilot, or ChatGPT.
  • Strong expertise in SLIs, SLOs, error budgets, incident command, blameless postmortems, and recurrence prevention for distributed systems.
  • Strong software engineering skills in Python, Go, or similar languages, with experience building automation and tools.
  • Deep experience with observability, infrastructure as code, Terraform, Kubernetes, AWS, and safe reversible deployments.

Nice to have

  • Experience in fintech, regulated environments, SOC 2, PCI, FinOps, or high-scale cost and performance tradeoffs.

Culture & Benefits

  • Hybrid work arrangement based in the Mountain View headquarters.
  • Equity and employee benefits included with the base salary.
  • Inclusive workplace focused on diversity, belonging, and representation of the community served.
  • AI-assisted operational workflows are treated as a standard engineering practice while retaining human accountability and judgment.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →