19 дней назад
Associate Staff Engineer (SRE)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Associate Staff Engineer (SRE) (AWS/Cloud/DevOps): Ensuring the reliability, availability, performance, scalability, and security of business-critical hospitality platforms with an accent on Kubernetes, Infrastructure as Code, observability, and production support. Focus on leading major-incident response, automating operational toil, validating disaster recovery, and supporting 24x7 production systems for reservations, payments, loyalty, and guest-facing applications.
Location: Shanghai, China; Service Region: China
Company
provides technology and engineering services, including support for global hospitality platforms and business-critical digital systems.
What you will do
- Ensure the availability, reliability, performance, scalability, security, and operational readiness of production services.
- Define and monitor SLIs, SLOs, SLAs, error budgets, availability, latency, and service health metrics.
- Act as a senior escalation point for P1/P2 incidents, leading troubleshooting, root cause analysis, post-incident reviews, and corrective actions.
- Operate workloads across AWS, Azure, and/or GCP, including Kubernetes platforms such as EKS, AKS, GKE, and OpenShift.
- Develop Infrastructure as Code, CI/CD pipelines, observability solutions, automation, self-healing, and operational documentation.
- Support disaster recovery, failover testing, capacity planning, releases, vulnerability remediation, and a global 24x7 on-call model.
Requirements
- Bachelor's degree in Computer Science, Engineering, Information Technology, or a related discipline.
- 7+ years of experience in Cloud, DevOps, Infrastructure, Production Engineering, or IT Operations, including 3+ years in SRE, DevOps, Cloud Operations, or Production Engineering.
- Strong hands-on AWS, Azure, or GCP experience, with strong Kubernetes, Docker, and container troubleshooting skills.
- Hands-on experience with Terraform or equivalent Infrastructure as Code technologies, CI/CD, deployment automation, Linux, production troubleshooting, and observability platforms.
- Experience with major incidents, RCA, problem management, distributed systems, APIs, databases, messaging platforms, networking, and ITIL-based processes.
- Strong written and verbal English communication skills.
Nice to have
- Experience in hospitality, travel, airline, e-commerce, financial services, or other 24x7 high-availability industries.
- Experience with high-volume transactional or reservation platforms, PCI DSS, GDPR, ISO 27001, DevSecOps, and vulnerability management.
- Knowledge of Kafka, Redis, API gateways, service mesh, event-driven architectures, FinOps, resilience testing, chaos engineering, or automated recovery.
- AWS, Azure, GCP, CKA, Terraform Associate, or ITIL certifications.
Culture & Benefits
- Work within a global 24x7 operating model alongside Development, Infrastructure, Security, Architecture, and Service Management teams.
- Support critical hospitality journeys including hotel search, booking, reservation changes, payments, check-in/check-out, loyalty transactions, and property-system integrations.
- Participate in production on-call and support activities as required.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →