Назад
Company hidden
6 дней назад

SRE Leader

Тип работы
fulltime
Грейд
lead
Английский
b2
Страна
Malaysia
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
SRE Leader (SRE/FinOps): Building company-wide reliability engineering, cloud cost governance, automated operations, and compliant multi-region infrastructure for a cryptocurrency exchange with an accent on SLOs, self-healing systems, FinOps, and financial-grade resilience. Focus on designing multi-cloud isolation, automating AIOps and disaster recovery, reducing operational toil, and cultivating a high-performing SRE organization.

Location: Kuala Lumpur, Malaysia

Company

hirify.global is a cryptocurrency exchange and digital financial platform serving users across more than 200 countries and regions, with products spanning trading, payments, wealth management, custody, institutional services, and Web3.

What you will do

  • Establish company-wide SLO, SLA, error budget, MTTD, and MTTR systems and use reliability metrics to guide investment and operational decisions.
  • Build self-healing, chaos engineering, canary release, automated rollback, incident management, and on-call improvement capabilities.
  • Develop data-driven FinOps and capacity planning systems, including cost attribution, optimization automation, and resource forecasting based on business metrics.
  • Implement GitOps and IaC practices, AIOps capabilities, automated anomaly detection, runbook execution, and self-service operations platforms.
  • Design financial-grade multi-account, multi-VPC, multi-region isolation and disaster recovery architectures across AWS and other cloud platforms.
  • Build and develop the SRE organization, competency model, knowledge-sharing practices, and succession coverage for critical systems.

Requirements

  • 10+ years of experience in infrastructure, operations, or SRE, including 5+ years leading teams of more than 10 SRE or infrastructure professionals.
  • Practical expertise in SLO/SLI, error budgets, toil management, capacity planning, and incident management.
  • Experience managing environments with annual cloud spending above $5 million and delivering data-driven FinOps optimization.
  • Large-scale IaC and automated operations experience with Terraform, Pulumi, or CloudFormation, plus AIOps implementation experience.
  • Financial-grade or compliance-focused infrastructure experience, including multi-account and multi-VPC isolation, multi-region deployments, data sovereignty, PCI-DSS, or SOC2.
  • AWS experience is required, along with at least one additional cloud platform such as Tencent Cloud, GCP, or Azure; ability to build automation systems in Go or Python.

Nice to have

  • SRE management experience in cryptocurrency exchanges, securities firms, or payment companies.
  • Large-scale Kubernetes operations involving 100+ clusters or 10,000+ nodes.
  • Experience with high-availability trading systems, internal FinOps platforms, chaos engineering, or SOC2, ISO 27001, and PCI-DSS audits.

Culture & Benefits

  • Engineering- and data-driven approach to reliability, cost efficiency, and scalable infrastructure design.
  • Support for professional development through a Study Growth Fund.
  • Regular team-building activities, workshops, and internal events.
  • International collaboration with colleagues from around the world.
  • Career advancement and internal mobility opportunities.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →