Back
Company hidden
2 months ago

Site Reliability Engineer (AWS/Kubernetes)

Work format
remote (only Europe)
Work type
fulltime
Grade
senior
English
b2
This vacancy is from Hirify.Global listVacancy from Hirify Global, list of international tech companies
Plus is required to make matches and apply

Match & Cover letter

Plus required for matching with this vacancy

Job description

Text:
/
TL;DR
Site Reliability Engineer (AWS/Kubernetes): Building and optimizing high-availability cloud infrastructure and Kubernetes platforms for a global sports betting platform with an accent on GitOps, observability, and resource optimization. Focus on automating environment provisioning, defining SLIs/SLOs, and ensuring system reliability under high traffic volumes.

Location: Must be based in Europe

Company

hirify.global is a remote-first sports betting and gaming company.

What you will do

  • Improve infrastructure and streamline deployment processes for new market expansions.
  • Optimize Kubernetes (EKS) stability and efficiency using GitOps-first practices with ArgoCD and Helm.
  • Manage cloud infrastructure via autoscaling, alerting pipelines, and Grafana dashboards.
  • Own weekend on-call operations, incident triaging, root cause analysis, and post-incident reviews.
  • Define and maintain SLIs/SLOs to drive reliability improvements and on-call prioritization.
  • Mentor junior team members and coordinate with security agencies for annual audits.

Requirements

  • 3+ years of experience in DevOps, SRE, or platform engineering.
  • Must be based in Europe.
  • Strong experience with AWS, Kubernetes (EKS), and Terraform.
  • Proficiency in scripting and automation using Bash, Python, or Golang.
  • Hands-on experience with observability stacks including Prometheus, Loki, Tempo, and OpenTelemetry.
  • Proven experience in on-call incident response and designing systems for high traffic environments.

Nice to have

  • Experience with Rust.
  • Familiarity with Grafana Faro or OpenTelemetry SDK instrumentation.
  • Knowledge of Cilium-based service mesh.
  • Expertise in JVM optimization.

Culture & Benefits

  • Remote-first culture with flexible core working hours (10am-3pm local time).
  • Competitive salary with quarterly individual performance bonuses.
  • 28 days of paid annual leave.
  • Top-of-the-line equipment provided.
  • Annual company retreats for internal networking.

Hiring process

  • Remote video screening with the Talent Acquisition Team.
  • Online technical assessment via Hackerrank.
  • Remote video interview with three team members.

Be careful: if the employer asks you to log into their system using iCloud/Google, send codes/passwords, or run code/software, don't do it - these are scammers. Always click "Report" or contact support. More in guide →