Назад
Company hidden
21 час назад

Senior Site Reliability Engineer (Payments Infrastructure)

Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/

TL;DR

Senior Site Reliability Engineer (Payments Infrastructure): Ensure the reliability, scalability, and operational excellence of a global payment platform with an accent on production observability, incident response, and cloud infrastructure reliability. Focus on operating distributed payment systems, leading SEV1/SEV2 response, improving MTTR, and building resilient services in PCI-DSS-regulated environments.

Location: Hybrid in Palo Alto, California, United States

Company

Develops and operates a global payment platform serving Europe, Asia, and North America.

What you will do

  • Own production observability, service-level management, and reliability across mission-critical payment processing systems.
  • Participate in a follow-the-sun on-call rotation and respond to production incidents across payment services, Kubernetes, databases, messaging systems, and cloud infrastructure.
  • Define SLOs, SLIs, error budgets, alerting standards, and operational readiness processes.
  • Drive reliability improvements through automation, capacity planning, performance optimization, and post-incident reviews.
  • Lead SEV1/SEV2 incident management and coordinate cross-functional response efforts.
  • Partner with engineering teams across the US and international hubs to improve resilience, security, and operational maturity.

Requirements

  • 5+ years of experience in SRE, platform engineering, DevOps, or cloud infrastructure roles supporting mission-critical production systems.
  • Hands-on experience with AWS, Kubernetes/EKS, Terraform, PostgreSQL, Redis, Kafka, Linux, networking, and modern observability platforms.
  • Strong understanding of distributed systems, cloud-native architectures, high availability, disaster recovery, capacity planning, and performance optimization.
  • Experience operating payment, banking, fintech, or other highly regulated systems with stringent security, compliance, and uptime requirements.
  • Strong knowledge of SRE principles, including SLOs, SLIs, error budgets, incident management, alert governance, and operational excellence.
  • Strong ownership, structured troubleshooting, data-driven decision-making, cross-functional communication, mentoring, and technical leadership skills.

Culture & Benefits

  • Competitive compensation aligned with California market standards.
  • Opportunity to lead work in a rapidly growing and innovative environment.
  • Collaborative and inclusive workplace where contributions are recognized.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →