Назад
Company hidden
3 дня назад

Senior Manager, Site Reliability Engineering (AI Ops)

222 000 - 300 500$
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Manager, Site Reliability Engineering (AI Ops) (AWS/Fintech): Leading a hands-on team that operates resilient AWS infrastructure and reliability systems for high-trust fintech services with an accent on 99.999% availability, incident management, and autonomous operations. Focus on designing AI-driven detection, diagnosis, and remediation, scaling multi-region infrastructure, and reducing operational toil and MTTR.

Location: Mountain View, California, United States

Base pay: $222,000–$300,500 per year, with potential bonus, equity rewards, and benefits.

Company

Financial technology platform providing TurboTax, Credit Karma, QuickBooks, and Mailchimp to approximately 100 million customers worldwide.

What you will do

  • Lead and develop a team of 10–15 systems and site reliability engineers.
  • Own operational excellence for fintech platform services and drive a 99.999% availability target.
  • Define and execute an AI Ops roadmap for autonomous detection, diagnosis, remediation, and toil reduction.
  • Lead incident command for high-severity incidents, improve postmortems, and strengthen on-call practices.
  • Build resilient AWS infrastructure with multi-AZ and multi-region failover, auto-remediation, and chaos engineering.
  • Partner with engineering, product, security, and compliance teams on reliability, observability, capacity, and infrastructure strategy.

Requirements

  • 8+ years of systems, site reliability, or infrastructure engineering experience, including 3+ years managing engineering teams.
  • Hands-on experience operating AWS infrastructure at scale, including EC2, EKS/ECS, VPC, RDS/DynamoDB, IAM, CloudWatch, and Auto Scaling.
  • Strong background in distributed systems, networking, Kubernetes, infrastructure as code, observability, SLOs, and incident management.
  • Experience delivering high availability for mission-critical customer-facing systems and communicating operational risk to senior leadership.
  • Bachelor’s degree in Computer Science, Engineering, or equivalent practical experience.
  • Experience with AI Ops or autonomous remediation, regulated fintech or payments environments, and chaos or resilience testing.

Culture & Benefits

  • Competitive compensation with performance-based rewards.
  • Potential cash bonus, equity rewards, and benefits according to applicable plans.
  • Focus on operational rigor, psychological safety, and continuous improvement.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →