8 дней назад
Reliability Engineer (Fintech)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Reliability Engineer (Fintech): Safeguarding production performance for a global trading platform with an accent on incident response, observability, triage automation, and service reliability. Focus on coordinating major incidents as Incident Commander, improving monitoring and alerting, automating operational workflows, and driving teams to resolve reliability risks.
Location: Hong Kong
Company
operates a global trading platform and the technology estate supporting it.
What you will do
- Automate repetitive triage workflows, including alert enrichment, routing, correlation, and operational tasks.
- Improve monitoring, alerting, and visibility across critical trading applications and their dependencies.
- Track reliability and availability, identify service degradation, and drive owning teams to resolve operational risks.
- Triage alerts and escalations, assess impact and ownership, and declare incidents when criteria are met.
- Act as Incident Commander, coordinating responders, stakeholders, mitigation, recovery, timelines, and status updates.
- Support post-incident reviews and recurring issue analysis while maintaining handovers across EMEA, AMER, and APAC.
Requirements
- Experience in production operations, SRE, NOC or command centre work, trading operations, or a similar first-line technical role.
- Experience running or coordinating major incidents and taking command of calls involving senior stakeholders.
- Strong triage, prioritisation, communication, judgment, and escalation skills.
- Solid Linux and networking fundamentals, with the ability to interpret alerts, logs, dashboards, and operational symptoms.
- Experience with incident and observability tooling such as PagerDuty, Jira Service Management, Grafana, Prometheus, and log search.
- Scripting and automation skills, preferably Python; familiarity with Bash, Go, Kubernetes, Docker, and GCP is useful.
Culture & Benefits
- Work alongside traders, developers, infrastructure, connectivity, data, and specialist teams.
- Operate under a global incident management framework with consistent standards and handover processes.
- Take ownership of operational standards and challenge weak ownership, poor alerts, missed SLAs, and ineffective runbooks.
- Support a latency-sensitive trading environment where rapid detection, containment, and recovery are critical.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →