4 Π΄Π½Ρ Π½Π°Π·Π°Π΄
Senior Site Reliability Engineer (SRE/AI)
ΠΡΡΡ & Π‘ΠΎΠΏΡΠΎΠ²ΠΎΠ΄
ΠΠ»Ρ ΠΌΡΡΡΠ° Ρ ΡΡΠΎΠΉ Π²Π°ΠΊΠ°Π½ΡΠΈΠ΅ΠΉ Π½ΡΠΆΠ΅Π½ Plus
ΠΠΏΠΈΡΠ°Π½ΠΈΠ΅ Π²Π°ΠΊΠ°Π½ΡΠΈΠΈ
Π’Π΅ΠΊΡΡ:
TL;DR
Senior Site Reliability Engineer (SRE/AI): Improving reliability for globally distributed trading and front-office systems with an accent on incident management, observability, automation, and reliability governance. Focus on designing self-healing mechanisms, managing SLOs and error budgets, reducing alert noise and operational toil, and leading cross-region incident response.
Location: Singapore, Singapore
Company
is a global investment banking, securities, and investment management firm founded in 1869, with offices worldwide.
What you will do
- Act as Incident Commander for high-severity incidents, coordinating escalation, resolver engagement, status updates, and post-incident reviews.
- Own cross-region handoffs, desk-readiness procedures, shift notes, and operational risk tracking.
- Automate repetitive operational work, codify remediations, implement self-healing mechanisms, and evaluate AI-assisted triage and anomaly detection.
- Improve observability, monitoring, alert quality, tracing, logging, and metrics across distributed systems.
- Define and govern SLOs, SLIs, error budgets, capacity standards, progressive delivery, and operational readiness reviews.
- Establish SRE KPIs and executive-ready reporting covering incident performance, change quality, toil reduction, and capacity headroom.
Requirements
- At least 5 years of experience in SRE, production operations, or reliability-focused engineering for high-availability customer-facing or trading systems.
- Proven Incident Commander experience with measurable improvements in escalation, communications, and MTTR.
- Strong knowledge of Linux, networking, distributed systems, and AWS, Azure, or GCP.
- Hands-on experience with Prometheus, Grafana, OpenTelemetry, ELK, PagerDuty or Opsgenie, and collaboration platforms.
- Experience with Terraform, CloudFormation, Ansible, and at least one modern programming language such as Go or Python.
- Excellent written and verbal communication, with the ability to explain technical incidents to executives and business stakeholders.
Nice to have
- Experience in front-office trading or other latency- and availability-sensitive environments.
- Kubernetes-based microservices, service meshes, multi-region architectures, or global standards harmonization.
- AI-assisted operations, status pages, customer-facing incident communications, ITIL-aligned processes, or ORR governance.
Culture & Benefits
- Training and development opportunities and firmwide professional networks.
- Benefits, wellness, personal finance offerings, and mindfulness programs.
- Commitment to diversity, inclusion, equal opportunity, and reasonable accommodations.
ΠΡΠ΄ΡΡΠ΅ ΠΎΡΡΠΎΡΠΎΠΆΠ½Ρ: Π΅ΡΠ»ΠΈ ΡΠ°Π±ΠΎΡΠΎΠ΄Π°ΡΠ΅Π»Ρ ΠΏΡΠΎΡΠΈΡ Π²ΠΎΠΉΡΠΈ Π² ΠΈΡ ΡΠΈΡΡΠ΅ΠΌΡ, ΠΈΡΠΏΠΎΠ»ΡΠ·ΡΡ iCloud/Google, ΠΏΡΠΈΡΠ»Π°ΡΡ ΠΊΠΎΠ΄/ΠΏΠ°ΡΠΎΠ»Ρ, Π·Π°ΠΏΡΡΡΠΈΡΡ ΠΊΠΎΠ΄/ΠΠ, Π½Π΅ Π΄Π΅Π»Π°ΠΉΡΠ΅ ΡΡΠΎΠ³ΠΎ - ΡΡΠΎ ΠΌΠΎΡΠ΅Π½Π½ΠΈΠΊΠΈ. ΠΠ±ΡΠ·Π°ΡΠ΅Π»ΡΠ½ΠΎ ΠΆΠΌΠΈΡΠ΅ "ΠΠΎΠΆΠ°Π»ΠΎΠ²Π°ΡΡΡΡ" ΠΈΠ»ΠΈ ΠΏΠΈΡΠΈΡΠ΅ Π² ΠΏΠΎΠ΄Π΄Π΅ΡΠΆΠΊΡ. ΠΠΎΠ΄ΡΠΎΠ±Π½Π΅Π΅ Π² Π³Π°ΠΉΠ΄Π΅ β
ΠΠΎΡ ΠΎΠΆΠΈΠ΅ Π²Π°ΠΊΠ°Π½ΡΠΈΠΈ
2 Π΄Π½Ρ Π½Π°Π·Π°Π΄
Senior Site Reliability Engineer (SRE)
130Β 000 - 175Β 000$
2 Π΄Π½Ρ Π½Π°Π·Π°Π΄
Customer Reliability Engineer (AWS)
5 Π΄Π½Π΅ΠΉ Π½Π°Π·Π°Π΄
Senior Software / Site Reliability Lead Engineer (AI)
142Β 696 - 158Β 303$
5 Π΄Π½Π΅ΠΉ Π½Π°Π·Π°Π΄
Senior Site Reliability Engineer
5 Π΄Π½Π΅ΠΉ Π½Π°Π·Π°Π΄
Senior Software Engineer (SRE)
3 Π΄Π½Ρ Π½Π°Π·Π°Π΄