Назад
Company hidden
4 дня назад

Senior Site Reliability Engineer (SRE/AI)

Π’ΠΈΠΏ Ρ€Π°Π±ΠΎΡ‚Ρ‹
fulltime
Π“Ρ€Π΅ΠΉΠ΄
senior
Английский
b2
Π‘Ρ‚Ρ€Π°Π½Π°
Singapore
Вакансия ΠΈΠ· списка Hirify.GlobalВакансия ΠΈΠ· Hirify Global, списка ΠΌΠ΅ΠΆΠ΄ΡƒΠ½Π°Ρ€ΠΎΠ΄Π½Ρ‹Ρ… tech-ΠΊΠΎΠΌΠΏΠ°Π½ΠΈΠΉ
Для мэтча ΠΈ ΠΎΡ‚ΠΊΠ»ΠΈΠΊΠ° Π½ΡƒΠΆΠ΅Π½ Plus

ΠœΡΡ‚Ρ‡ & Π‘ΠΎΠΏΡ€ΠΎΠ²ΠΎΠ΄

Для мэтча с этой вакансиСй Π½ΡƒΠΆΠ΅Π½ Plus

ОписаниС вакансии

ВСкст:
/
TL;DR
Senior Site Reliability Engineer (SRE/AI): Improving reliability for globally distributed trading and front-office systems with an accent on incident management, observability, automation, and reliability governance. Focus on designing self-healing mechanisms, managing SLOs and error budgets, reducing alert noise and operational toil, and leading cross-region incident response.

Location: Singapore, Singapore

Company

hirify.global is a global investment banking, securities, and investment management firm founded in 1869, with offices worldwide.

What you will do

  • Act as Incident Commander for high-severity incidents, coordinating escalation, resolver engagement, status updates, and post-incident reviews.
  • Own cross-region handoffs, desk-readiness procedures, shift notes, and operational risk tracking.
  • Automate repetitive operational work, codify remediations, implement self-healing mechanisms, and evaluate AI-assisted triage and anomaly detection.
  • Improve observability, monitoring, alert quality, tracing, logging, and metrics across distributed systems.
  • Define and govern SLOs, SLIs, error budgets, capacity standards, progressive delivery, and operational readiness reviews.
  • Establish SRE KPIs and executive-ready reporting covering incident performance, change quality, toil reduction, and capacity headroom.

Requirements

  • At least 5 years of experience in SRE, production operations, or reliability-focused engineering for high-availability customer-facing or trading systems.
  • Proven Incident Commander experience with measurable improvements in escalation, communications, and MTTR.
  • Strong knowledge of Linux, networking, distributed systems, and AWS, Azure, or GCP.
  • Hands-on experience with Prometheus, Grafana, OpenTelemetry, ELK, PagerDuty or Opsgenie, and collaboration platforms.
  • Experience with Terraform, CloudFormation, Ansible, and at least one modern programming language such as Go or Python.
  • Excellent written and verbal communication, with the ability to explain technical incidents to executives and business stakeholders.

Nice to have

  • Experience in front-office trading or other latency- and availability-sensitive environments.
  • Kubernetes-based microservices, service meshes, multi-region architectures, or global standards harmonization.
  • AI-assisted operations, status pages, customer-facing incident communications, ITIL-aligned processes, or ORR governance.

Culture & Benefits

  • Training and development opportunities and firmwide professional networks.
  • Benefits, wellness, personal finance offerings, and mindfulness programs.
  • Commitment to diversity, inclusion, equal opportunity, and reasonable accommodations.

Π‘ΡƒΠ΄ΡŒΡ‚Π΅ остороТны: Ссли Ρ€Π°Π±ΠΎΡ‚ΠΎΠ΄Π°Ρ‚Π΅Π»ΡŒ просит Π²ΠΎΠΉΡ‚ΠΈ Π² ΠΈΡ… систСму, ΠΈΡΠΏΠΎΠ»ΡŒΠ·ΡƒΡ iCloud/Google, ΠΏΡ€ΠΈΡΠ»Π°Ρ‚ΡŒ ΠΊΠΎΠ΄/ΠΏΠ°Ρ€ΠΎΠ»ΡŒ, Π·Π°ΠΏΡƒΡΡ‚ΠΈΡ‚ΡŒ ΠΊΠΎΠ΄/ПО, Π½Π΅ Π΄Π΅Π»Π°ΠΉΡ‚Π΅ этого - это мошСнники. ΠžΠ±ΡΠ·Π°Ρ‚Π΅Π»ΡŒΠ½ΠΎ ΠΆΠΌΠΈΡ‚Π΅ "ΠŸΠΎΠΆΠ°Π»ΠΎΠ²Π°Ρ‚ΡŒΡΡ" ΠΈΠ»ΠΈ ΠΏΠΈΡˆΠΈΡ‚Π΅ Π² ΠΏΠΎΠ΄Π΄Π΅Ρ€ΠΆΠΊΡƒ. ΠŸΠΎΠ΄Ρ€ΠΎΠ±Π½Π΅Π΅ Π² Π³Π°ΠΉΠ΄Π΅ β†’