Senior Site Reliability Engineer (Cloud)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
TL;DR
Senior Site Reliability Engineer (Cloud): Building and maintaining a highly available cloud platform for a sports ticketing and streaming service with an accent on observability, automation, and system reliability. Focus on designing monitoring capabilities, improving CI/CD workflows, and translating Critical User Journeys into SLA/SLO objectives.
Location: Almaty, Astana, Belgrade, Cluj-Napoca, Dnipro, Kharkiv, Krakow, Kyiv, Larnaca, Lodz, Lublin, Lviv, Odesa, Riga, Sofia, Tbilisi, Varna, Warsaw, Wroclaw, Yerevan; Remote in Bulgaria, Georgia, Kazakhstan, Poland. A 4-hour overlap with Eastern Standard Time (EST) is required.
Company
is a global IT service provider specializing in the development of scalable and high-performing software solutions for diverse industries.
What you will do
- Enhance platform observability by designing metrics, alerts, and dashboards to reduce incident resolution time.
- Develop automation, operational tooling, and monitoring solutions to increase service reliability and uptime.
- Collaborate with software development and QA teams to embed reliability best practices into software delivery and release processes.
- Drive operational excellence through preventive measures and facilitating blameless post-incident reviews.
- Participate in an on-call rotation to ensure timely resolution of production incidents affecting critical services.
Requirements
- Strong hands-on experience with Python for scripting and automation.
- Proficiency in at least one language: Java, C++, or Go.
- Solid knowledge of Linux, cloud platforms (AWS, GCP, or Azure), and containerized infrastructure (Docker, Kubernetes, Terraform).
- Experience designing and maintaining CI/CD pipelines and implementing automated testing.
- Practical experience with monitoring platforms such as Prometheus, Grafana, ELK, or Datadog.
- Must be able to ensure a 4-hour overlap with EST working hours.
Nice to have
- Experience with end-to-end and integration tests for microservices-based systems.
- Knowledge of chaos engineering, performance optimization, or capacity management.
- Experience contributing to internal developer platforms or reliability initiatives.
- Understanding of production security, compliance, and change management.
Culture & Benefits
- Health insurance for employees and their loved ones.
- Vacation and sick pay according to the laws of the employee's country.
- Paid time off for state holidays regardless of the client's schedule.
- Support for professional growth, including coverage for IT certifications and access to learning platforms.
- A supportive environment with corporate parties and flexible work options.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →