2 Π΄Π½Ρ Π½Π°Π·Π°Π΄
Site Reliability Engineer III (GCP)
ΠΡΡΡ & Π‘ΠΎΠΏΡΠΎΠ²ΠΎΠ΄
ΠΠ»Ρ ΠΌΡΡΡΠ° Ρ ΡΡΠΎΠΉ Π²Π°ΠΊΠ°Π½ΡΠΈΠ΅ΠΉ Π½ΡΠΆΠ΅Π½ Plus
ΠΠΏΠΈΡΠ°Π½ΠΈΠ΅ Π²Π°ΠΊΠ°Π½ΡΠΈΠΈ
Π’Π΅ΠΊΡΡ:
TL;DR
Site Reliability Engineer III (GCP): Building and operating resilient, automated GCP infrastructure and middleware platforms for clearing, risk, and derivatives applications with an accent on ultra-low-latency performance, high concurrency, and observability. Focus on designing reliability tooling, leading production incident recovery, reducing operational toil through automation, and testing disaster recovery and system resiliency.
Location: Belfast, United Kingdom; hybrid working
Company
operates a leading global derivatives marketplace and builds technology for clearing, risk, and derivatives applications.
What you will do
- Architect, operate, and migrate messaging, service discovery, and data distribution platforms to Google Cloud.
- Manage cluster lifecycles, data replication, RBAC, and workload placement across middleware platforms.
- Design and maintain observability systems using OpenTelemetry, Splunk, Prometheus, and Grafana, including metrics, logs, alerts, SLIs, and SLOs.
- Respond to production incidents, lead post-mortems, and drive rapid system recovery.
- Reduce operational toil through automation, code, disaster recovery strategies, and continuous resiliency testing.
- Lead technical discussions, collaborate across functions, and mentor junior SRE colleagues.
Requirements
- Programming or scripting experience with Python, Go, Java, or Bash.
- Proficiency with Linux, distributed systems, containerization, Kubernetes/GKE, and GCP/GCE.
- Understanding of CI/CD and infrastructure-as-code tools such as Terraform, Ansible, or Kubernetes Config Connector.
- Knowledge of TCP/IP, UDP, HTTP, DNS, load balancing, and messaging protocols.
- Ability to troubleshoot complex system behavior and communicate technical requirements across teams.
- Experience applying generative AI and agents such as Gemini to platform operations.
Nice to have
- Hands-on experience with OpenTelemetry, Splunk, Prometheus, and Grafana.
- Experience with Agile development practices and software development lifecycles.
- GCP Professional Cloud Architect, CKA, or CKAD certification.
- Experience in financial markets or other regulated, ultra-low-latency, high-concurrency environments.
Culture & Benefits
- Code-first engineering culture focused on systematic automation.
- Bonus, equity, and employee stock purchase programmes.
- Private medical and dental coverage, mental health benefits, pension, income protection, and life assurance.
- Family leave, education assistance, and ongoing development and certification training.
- Cycle-to-work scheme, EV car benefit scheme, and gym membership.
- Hybrid working arrangement.
ΠΡΠ΄ΡΡΠ΅ ΠΎΡΡΠΎΡΠΎΠΆΠ½Ρ: Π΅ΡΠ»ΠΈ ΡΠ°Π±ΠΎΡΠΎΠ΄Π°ΡΠ΅Π»Ρ ΠΏΡΠΎΡΠΈΡ Π²ΠΎΠΉΡΠΈ Π² ΠΈΡ ΡΠΈΡΡΠ΅ΠΌΡ, ΠΈΡΠΏΠΎΠ»ΡΠ·ΡΡ iCloud/Google, ΠΏΡΠΈΡΠ»Π°ΡΡ ΠΊΠΎΠ΄/ΠΏΠ°ΡΠΎΠ»Ρ, Π·Π°ΠΏΡΡΡΠΈΡΡ ΠΊΠΎΠ΄/ΠΠ, Π½Π΅ Π΄Π΅Π»Π°ΠΉΡΠ΅ ΡΡΠΎΠ³ΠΎ - ΡΡΠΎ ΠΌΠΎΡΠ΅Π½Π½ΠΈΠΊΠΈ. ΠΠ±ΡΠ·Π°ΡΠ΅Π»ΡΠ½ΠΎ ΠΆΠΌΠΈΡΠ΅ "ΠΠΎΠΆΠ°Π»ΠΎΠ²Π°ΡΡΡΡ" ΠΈΠ»ΠΈ ΠΏΠΈΡΠΈΡΠ΅ Π² ΠΏΠΎΠ΄Π΄Π΅ΡΠΆΠΊΡ. ΠΠΎΠ΄ΡΠΎΠ±Π½Π΅Π΅ Π² Π³Π°ΠΉΠ΄Π΅ β
ΠΠΎΡ ΠΎΠΆΠΈΠ΅ Π²Π°ΠΊΠ°Π½ΡΠΈΠΈ
3 Π΄Π½Ρ Π½Π°Π·Π°Π΄
Senior Site Reliability Engineer (Cloud Security)
72Β 702 - 80Β 780GBP
3 Π΄Π½Ρ Π½Π°Π·Π°Π΄
Site Reliability Engineer (Cloud Banking)
2 Π΄Π½Ρ Π½Π°Π·Π°Π΄
Site Reliability Engineer (AWS)
55Β 000 - 63Β 000GBP
1 Π΄Π΅Π½Ρ Π½Π°Π·Π°Π΄
Site Reliability Engineer (Kubernetes)
84Β 051 - 93Β 390GBP
3 Π΄Π½Ρ Π½Π°Π·Π°Π΄
Site Reliability Engineer (Cybersecurity)
48Β 987 - 54Β 430GBP
3 Π΄Π½Ρ Π½Π°Π·Π°Π΄
Site Reliability Engineer
3Β 500 - 4Β 500β¬