2 месяца назад
Senior Principal Site Reliability Engineer (Web3)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Principal Site Reliability Engineer (Web3): Building an enterprise-grade chaos engineering platform for multi-cluster, multi-region Kubernetes and EC2 environments with an accent on fault injection, production safety, and resilience validation. Focus on designing complex experiment orchestration, integrating monitoring and SLO systems into closed-loop validation, and mentoring engineers in production stability practices.
Location: Hong Kong SAR
Company
is a cryptocurrency exchange and digital financial platform serving users across more than 200 countries and regions, with products spanning trading, payments, wealth management, custody, institutional services, and Web3.
What you will do
- Design and build an enterprise-grade chaos engineering platform for multi-cluster Kubernetes and EC2 environments across multiple regions and environments.
- Develop fault injection capabilities for pods, nodes, availability zones, network conditions, dependencies, and complex combined failure scenarios.
- Implement production safety controls including blast-radius management, kill switches, automatic rollback, traffic isolation, and real-time impact monitoring.
- Integrate experiments with monitoring, alerting, and SLO systems to create automated fault-injection validation loops.
- Define approval standards, run routine resilience experiments and disaster-recovery drills, and establish resilience scoring.
- Select technology foundations, create playbooks, and mentor 2–3 engineers in chaos engineering.
Requirements
- 8+ years of backend or infrastructure engineering experience, including 3+ years in chaos engineering or stability engineering.
- Hands-on experience with large-scale fault injection in production and a deep understanding of production safety constraints.
- Expertise in Kubernetes fault injection using Chaos Mesh, Litmus, or custom solutions, including CRD and Operator development.
- Proficiency in at least one backend language, preferably Go, and the ability to design platform-level architectures.
- Strong understanding of distributed-system failure modes, including network partitions, split-brain, cascading failures, and data inconsistency.
- Experience with observability technologies such as Prometheus, Grafana, Thanos, or OpenTelemetry, plus strong technical documentation and solution-design skills.
Nice to have
- Financial or trading-system stability experience, including transaction consistency and fund-safety constraints.
- Experience with SLO and error-budget frameworks, automated fault recovery, AWS infrastructure, or established chaos engineering tools.
- Open-source contributions to Chaos Mesh, Litmus, or similar projects.
Culture & Benefits
- Study Growth Fund for professional development and continuous learning.
- Internal team-building activities, workshops, and collaboration events.
- International collaboration with colleagues from around the world.
- Career advancement and internal mobility opportunities.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →