6 дней назад
SRE Leader
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
SRE Leader (SRE/FinOps): Building company-wide reliability engineering, cloud cost governance, automated operations, and compliant multi-region infrastructure for a cryptocurrency exchange with an accent on SLOs, self-healing systems, FinOps, and financial-grade resilience. Focus on designing multi-cloud isolation, automating AIOps and disaster recovery, reducing operational toil, and cultivating a high-performing SRE organization.
Location: Kuala Lumpur, Malaysia
Company
is a cryptocurrency exchange and digital financial platform serving users across more than 200 countries and regions, with products spanning trading, payments, wealth management, custody, institutional services, and Web3.
What you will do
- Establish company-wide SLO, SLA, error budget, MTTD, and MTTR systems and use reliability metrics to guide investment and operational decisions.
- Build self-healing, chaos engineering, canary release, automated rollback, incident management, and on-call improvement capabilities.
- Develop data-driven FinOps and capacity planning systems, including cost attribution, optimization automation, and resource forecasting based on business metrics.
- Implement GitOps and IaC practices, AIOps capabilities, automated anomaly detection, runbook execution, and self-service operations platforms.
- Design financial-grade multi-account, multi-VPC, multi-region isolation and disaster recovery architectures across AWS and other cloud platforms.
- Build and develop the SRE organization, competency model, knowledge-sharing practices, and succession coverage for critical systems.
Requirements
- 10+ years of experience in infrastructure, operations, or SRE, including 5+ years leading teams of more than 10 SRE or infrastructure professionals.
- Practical expertise in SLO/SLI, error budgets, toil management, capacity planning, and incident management.
- Experience managing environments with annual cloud spending above $5 million and delivering data-driven FinOps optimization.
- Large-scale IaC and automated operations experience with Terraform, Pulumi, or CloudFormation, plus AIOps implementation experience.
- Financial-grade or compliance-focused infrastructure experience, including multi-account and multi-VPC isolation, multi-region deployments, data sovereignty, PCI-DSS, or SOC2.
- AWS experience is required, along with at least one additional cloud platform such as Tencent Cloud, GCP, or Azure; ability to build automation systems in Go or Python.
Nice to have
- SRE management experience in cryptocurrency exchanges, securities firms, or payment companies.
- Large-scale Kubernetes operations involving 100+ clusters or 10,000+ nodes.
- Experience with high-availability trading systems, internal FinOps platforms, chaos engineering, or SOC2, ISO 27001, and PCI-DSS audits.
Culture & Benefits
- Engineering- and data-driven approach to reliability, cost efficiency, and scalable infrastructure design.
- Support for professional development through a Study Growth Fund.
- Regular team-building activities, workshops, and internal events.
- International collaboration with colleagues from around the world.
- Career advancement and internal mobility opportunities.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
6 дней назад
SRE Engineering Manager (Cybersecurity)
7 дней назад
Engineering Leader (DevOps/SRE)
250 000 - 300 000$
7 дней назад
Tech Lead Manager (Go)
18 часов назад
Head of Platform Engineering
7 дней назад
Principal Tech Lead Manager (Data Platform & Reliability Engineering)
242 693 - 275 400$
6 дней назад