Senior Site Reliability Engineer (Payments Infrastructure)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
TL;DR
Senior Site Reliability Engineer (Payments Infrastructure): Ensure the reliability, scalability, and operational excellence of a global payment platform with an accent on production observability, incident response, and cloud infrastructure reliability. Focus on operating distributed payment systems, leading SEV1/SEV2 response, improving MTTR, and building resilient services in PCI-DSS-regulated environments.
Location: Hybrid in Palo Alto, California, United States
Company
Develops and operates a global payment platform serving Europe, Asia, and North America.
What you will do
- Own production observability, service-level management, and reliability across mission-critical payment processing systems.
- Participate in a follow-the-sun on-call rotation and respond to production incidents across payment services, Kubernetes, databases, messaging systems, and cloud infrastructure.
- Define SLOs, SLIs, error budgets, alerting standards, and operational readiness processes.
- Drive reliability improvements through automation, capacity planning, performance optimization, and post-incident reviews.
- Lead SEV1/SEV2 incident management and coordinate cross-functional response efforts.
- Partner with engineering teams across the US and international hubs to improve resilience, security, and operational maturity.
Requirements
- 5+ years of experience in SRE, platform engineering, DevOps, or cloud infrastructure roles supporting mission-critical production systems.
- Hands-on experience with AWS, Kubernetes/EKS, Terraform, PostgreSQL, Redis, Kafka, Linux, networking, and modern observability platforms.
- Strong understanding of distributed systems, cloud-native architectures, high availability, disaster recovery, capacity planning, and performance optimization.
- Experience operating payment, banking, fintech, or other highly regulated systems with stringent security, compliance, and uptime requirements.
- Strong knowledge of SRE principles, including SLOs, SLIs, error budgets, incident management, alert governance, and operational excellence.
- Strong ownership, structured troubleshooting, data-driven decision-making, cross-functional communication, mentoring, and technical leadership skills.
Culture & Benefits
- Competitive compensation aligned with California market standards.
- Opportunity to lead work in a rapidly growing and innovative environment.
- Collaborative and inclusive workplace where contributions are recognized.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →