12 часов назад
Site Reliability Engineer (iGaming)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer (iGaming): Building and operating enterprise observability, disaster recovery, and business continuity capabilities for high-performance gaming and betting platforms with an accent on AWS cloud infrastructure, monitoring, and system resilience. Focus on designing chaos and disaster recovery tests, optimizing systems for peak sporting-event loads, and improving incident response across regulated multi-market operations.
Location: Sofia, Bulgaria; hybrid work combining home and office attendance.
Company
operates sports betting, iGaming, and entertainment products across more than 100 global markets through over 8,000 colleagues and 28 offices.
What you will do
- Maintain 99.9%+ uptime for enterprise observability platforms supporting systems used by millions of concurrent users.
- Design and operate monitoring, alerting, dashboards, APM, telemetry, and data-ingestion solutions using tools such as Grafana, Splunk, CloudWatch, Prometheus, and ELK.
- Define SLOs and SLIs, perform capacity planning, and optimize performance for peak loads during major sporting events.
- Lead incident response improvements, post-incident reviews, blameless post-mortems, runbooks, and preventative reliability work with development and service management teams.
- Own chaos-testing frameworks and conduct disaster recovery fire drills in isolated environments to validate resilience and recovery procedures.
- Participate in architecture reviews, improve deployment reliability, mentor junior engineers, and document operational procedures and system architecture.
Requirements
- Extensive experience with enterprise-scale monitoring and observability tools, including Prometheus, Grafana, ELK, or similar solutions.
- Strong experience with AWS, Azure, or Google Cloud Platform and cloud system architecture.
- Experience applying reliability engineering practices in 24/7/365 production operations and in highly regulated, security-compliant environments.
- Strong scripting or programming skills in Python, Go, Bash, TypeScript, or Terraform, including infrastructure as code and automation.
- Experience with CI/CD, SQL and NoSQL databases, Docker, and Kubernetes.
- Ability to collaborate across functions, communicate during escalations, and produce clear technical documentation and runbooks.
Nice to have
- Previous software engineering experience.
- AWS certifications.
- Experience in gaming, financial services, healthcare, or another highly regulated industry.
Culture & Benefits
- Hybrid work with home and modern office options in Sofia.
- Annual discretionary bonus and 30 days of paid leave.
- Health and dental insurance, life insurance, disability coverage, and a wellbeing fund.
- Continuous learning support for certifications and career development.
- Paid maternity and paternity leave, sports card membership, monthly food vouchers, and service discounts.
- Inclusive workplace focused on diversity, accessibility, and responsible gaming.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
18 часов назад
Senior Site Reliability Engineer
7 дней назад
GCP Platform Engineer (Google Cloud)
11 часов назад
Senior Site Reliability Engineer (Azure Platform)
7 дней назад
Senior DevOps Engineer (Fintech)
1 день назад
Site Reliability Engineer (Prometheus/Grafana)
1 час назад
DevOps Deployment Engineer (AWS)
3 500 - 4 600€