3 дня назад
Senior Manager, Site Reliability Engineering (AI Ops)
222 000 - 300 500$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Manager, Site Reliability Engineering (AI Ops) (AWS/Fintech): Leading a hands-on team that operates resilient AWS infrastructure and reliability systems for high-trust fintech services with an accent on 99.999% availability, incident management, and autonomous operations. Focus on designing AI-driven detection, diagnosis, and remediation, scaling multi-region infrastructure, and reducing operational toil and MTTR.
Location: Mountain View, California, United States
Base pay: $222,000–$300,500 per year, with potential bonus, equity rewards, and benefits.
Company
Financial technology platform providing TurboTax, Credit Karma, QuickBooks, and Mailchimp to approximately 100 million customers worldwide.
What you will do
- Lead and develop a team of 10–15 systems and site reliability engineers.
- Own operational excellence for fintech platform services and drive a 99.999% availability target.
- Define and execute an AI Ops roadmap for autonomous detection, diagnosis, remediation, and toil reduction.
- Lead incident command for high-severity incidents, improve postmortems, and strengthen on-call practices.
- Build resilient AWS infrastructure with multi-AZ and multi-region failover, auto-remediation, and chaos engineering.
- Partner with engineering, product, security, and compliance teams on reliability, observability, capacity, and infrastructure strategy.
Requirements
- 8+ years of systems, site reliability, or infrastructure engineering experience, including 3+ years managing engineering teams.
- Hands-on experience operating AWS infrastructure at scale, including EC2, EKS/ECS, VPC, RDS/DynamoDB, IAM, CloudWatch, and Auto Scaling.
- Strong background in distributed systems, networking, Kubernetes, infrastructure as code, observability, SLOs, and incident management.
- Experience delivering high availability for mission-critical customer-facing systems and communicating operational risk to senior leadership.
- Bachelor’s degree in Computer Science, Engineering, or equivalent practical experience.
- Experience with AI Ops or autonomous remediation, regulated fintech or payments environments, and chaos or resilience testing.
Culture & Benefits
- Competitive compensation with performance-based rewards.
- Potential cash bonus, equity rewards, and benefits according to applicable plans.
- Focus on operational rigor, psychological safety, and continuous improvement.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
7 дней назад
Senior Engineering Manager, Infrastructure (AI)
230 000 - 280 000$
5 дней назад
Principal Tech Lead Manager (Data Platform & Reliability Engineering)
242 693 - 275 400$
Okta
5 дней назад
Senior Manager, Site Reliability Engineering - Infrastructure Platform (AWS)
232 000 - 319 000$
4 дня назад
Manager, Platform Engineering (AI)
121 000 - 178 000$
4 дня назад
Engineering Manager, Platform (AI)
217 000 - 260 000$
Affirm
3 дня назад
Manager, Software Engineering (Reliability Platform)
230 000 - 290 000$