3 дня назад
Staff Site Reliability Engineer (AI)
252 000 - 308 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Staff Site Reliability Engineer (AI) (AWS/Kubernetes): Defining reliability standards and building AI-assisted operational workflows for critical financial services with an accent on incident response, observability, resilience, and production readiness. Focus on designing AI-driven alert triage and runbook automation, leading high-severity incidents, reducing on-call toil, and guiding large-scale AWS architecture across EKS, Kafka, DynamoDB, RDS, and SQS.
Location: Mountain View, US; hybrid position requiring in-office work 2 days a week
Salary: $252,000–$308,000 base salary annually, plus equity and benefits
Company
builds earned wage access products that provide real-time financial flexibility, enabling members to access, spend, save, and grow their gs.
What you will do
- Define reliability standards across critical services, including SLIs, SLOs, error budgets, observability, production readiness, incident response, and resilience patterns.
- Build AI-assisted workflows for alert correlation, incident triage, root-cause exploration, runbook retrieval, postmortem drafting, and corrective-action tracking.
- Lead high-severity incidents as Incident Commander and improve detection, response quality, communication, and recurrence prevention.
- Develop tools that reduce on-call toil by gathering context from Datadog, CloudWatch, incident.io, Slack, runbooks, deployment history, and service metadata.
- Guide resilient service architecture across AWS, including graceful degradation, failure isolation, capacity planning, and operational safety for EKS, Kafka, DynamoDB, RDS, and SQS.
- Mentor engineers and influence product engineering, infrastructure, security, and leadership teams through design reviews, incident reviews, documentation, and reusable reliability practices.
Requirements
- 7+ years of experience in SRE, software engineering, or infrastructure engineering with increasing scope and cross-organizational influence.
- Demonstrated success improving reliability and operational excellence at scale using KPIs such as MTTR, MTTD, alert quality, incident recurrence, SLO attainment, on-call health, and corrective-action completion.
- Experience applying AI or LLMs to engineering and operational workflows, plus practical use of AI-assisted development tools such as Cursor, Claude Code, Copilot, or ChatGPT.
- Strong expertise in SLIs, SLOs, error budgets, incident command, blameless postmortems, and recurrence prevention for distributed systems.
- Strong software engineering skills in Python, Go, or similar languages, with experience building automation and tools.
- Deep experience with observability, infrastructure as code, Terraform, Kubernetes, AWS, and safe reversible deployments.
Nice to have
- Experience in fintech, regulated environments, SOC 2, PCI, FinOps, or high-scale cost and performance tradeoffs.
Culture & Benefits
- Hybrid work arrangement based in the Mountain View headquarters.
- Equity and employee benefits included with the base salary.
- Inclusive workplace focused on diversity, belonging, and representation of the community served.
- AI-assisted operational workflows are treated as a standard engineering practice while retaining human accountability and judgment.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
5 дней назад
Senior Site Reliability Engineer (AI)
191 000 - 226 000$
Okta
5 дней назад
Staff Site Reliability Engineer (Splunk)
194 000 - 267 000$
Okta
5 дней назад
Staff Site Reliability Engineer, Federal (TS/SCI)
174 000 - 238 000$
8 дней назад
Senior Manager, Site Reliability Engineering (AI Ops)
222 000 - 300 500$
Nscale
5 дней назад
Senior Site Reliability Engineer (AI Infrastructure Operations)
170 000 - 265 000$
Anthropic
9 дней назад
Staff+ Site Reliability Engineer (Safeguards ML Infra)
320 000 - 485 000$