Назад
Company hidden
12 часов назад

Site Reliability Engineer (iGaming)

Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Bulgaria
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Site Reliability Engineer (iGaming): Building and operating enterprise observability, disaster recovery, and business continuity capabilities for high-performance gaming and betting platforms with an accent on AWS cloud infrastructure, monitoring, and system resilience. Focus on designing chaos and disaster recovery tests, optimizing systems for peak sporting-event loads, and improving incident response across regulated multi-market operations.

Location: Sofia, Bulgaria; hybrid work combining home and office attendance.

Company

hirify.global operates sports betting, iGaming, and entertainment products across more than 100 global markets through over 8,000 colleagues and 28 offices.

What you will do

  • Maintain 99.9%+ uptime for enterprise observability platforms supporting systems used by millions of concurrent users.
  • Design and operate monitoring, alerting, dashboards, APM, telemetry, and data-ingestion solutions using tools such as Grafana, Splunk, CloudWatch, Prometheus, and ELK.
  • Define SLOs and SLIs, perform capacity planning, and optimize performance for peak loads during major sporting events.
  • Lead incident response improvements, post-incident reviews, blameless post-mortems, runbooks, and preventative reliability work with development and service management teams.
  • Own chaos-testing frameworks and conduct disaster recovery fire drills in isolated environments to validate resilience and recovery procedures.
  • Participate in architecture reviews, improve deployment reliability, mentor junior engineers, and document operational procedures and system architecture.

Requirements

  • Extensive experience with enterprise-scale monitoring and observability tools, including Prometheus, Grafana, ELK, or similar solutions.
  • Strong experience with AWS, Azure, or Google Cloud Platform and cloud system architecture.
  • Experience applying reliability engineering practices in 24/7/365 production operations and in highly regulated, security-compliant environments.
  • Strong scripting or programming skills in Python, Go, Bash, TypeScript, or Terraform, including infrastructure as code and automation.
  • Experience with CI/CD, SQL and NoSQL databases, Docker, and Kubernetes.
  • Ability to collaborate across functions, communicate during escalations, and produce clear technical documentation and runbooks.

Nice to have

  • Previous software engineering experience.
  • AWS certifications.
  • Experience in gaming, financial services, healthcare, or another highly regulated industry.

Culture & Benefits

  • Hybrid work with home and modern office options in Sofia.
  • Annual discretionary bonus and 30 days of paid leave.
  • Health and dental insurance, life insurance, disability coverage, and a wellbeing fund.
  • Continuous learning support for certifications and career development.
  • Paid maternity and paternity leave, sports card membership, monthly food vouchers, and service discounts.
  • Inclusive workplace focused on diversity, accessibility, and responsible gaming.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →