5 дней назад
Senior Site Reliability Engineer (Fintech)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Site Reliability Engineer (AWS/Kubernetes/Fintech): Driving reliability, availability, scalability, and operational excellence for a global payment platform with an accent on production observability, incident response, and cloud infrastructure reliability. Focus on leading SEV1/SEV2 response, designing SLOs and error budgets, improving distributed systems resilience, and operating PCI-DSS-regulated payment services.
Location: On-site in Shenzhen, China; candidates must be based in Hong Kong or Shenzhen
Company
operates a global payment platform serving payment processing needs across Europe, Asia, and North America.
What you will do
- Lead production incident management and participate in a follow-the-sun on-call rotation, including SEV1 and SEV2 response.
- Diagnose, mitigate, and coordinate resolution of incidents across payment services, Kubernetes platforms, databases, messaging systems, and cloud infrastructure.
- Define and maintain SLOs, SLIs, error budgets, alerting standards, and operational readiness processes.
- Improve reliability through infrastructure automation, observability, capacity planning, performance tuning, and root-cause analysis.
- Strengthen resilience, security, and operational maturity in PCI-DSS-regulated payment environments.
- Mentor engineers and promote resilience-by-design practices and blameless incident management.
Requirements
- 8+ years of hands-on experience in SRE, platform engineering, DevOps, or cloud infrastructure roles supporting high-availability production systems.
- Strong expertise in AWS, Kubernetes/EKS, Terraform, PostgreSQL, Redis, Kafka, Linux, networking, and observability platforms such as Datadog, Prometheus, or Grafana.
- Deep understanding of distributed systems, high availability, disaster recovery, capacity planning, and microservices orchestration.
- Experience operating payment, banking, fintech, or other highly regulated systems with PCI-DSS, security, and uptime requirements.
- Advanced knowledge of SRE practices, including SLO/SLI design, error budgets, alert governance, and toil reduction.
- Must be based in Hong Kong or Shenzhen and have excellent written and spoken English.
Culture & Benefits
- Competitive package.
- Dynamic and innovative working environment.
- Collaborative and inclusive team culture where contributions are recognized.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →