обновлено 2 часа назад
Manager, Site Reliability Engineering
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Manager, Site Reliability Engineering (Cloud/DevOps): Leading SRE teams in maintaining reliable, scalable, and high-performance systems and services with an accent on multi-cloud infrastructure, automation, observability, and incident management. Focus on designing high-availability and disaster-recovery strategies, improving infrastructure efficiency, and coordinating complex reliability work across software development and operations teams.
Location: Sofia, Bulgaria; hybrid
Company
is a travel technology company developing AI-driven airline retailing and intelligent offer and revenue optimization solutions.
What you will do
- Lead and mentor the Site Reliability Engineering team while promoting reliability, accountability, and continuous improvement.
- Develop strategies for multi-cloud reliability, monitoring, incident response, and system performance.
- Drive automation for deployment processes, infrastructure as code, configuration management, and operational efficiency.
- Manage observability for logging, metrics, and alerting, and establish SLOs, SLIs, and SLAs.
- Oversee root cause analysis and post-mortem processes, and promote SRE and DevOps best practices.
- Ensure high availability and disaster recovery, while optimizing cloud infrastructure costs.
Requirements
- Bachelor’s or master’s degree in computer science, engineering, or a related field.
- 7+ years of experience in software engineering, SRE, or DevOps, including at least 3 years in a managerial or leadership role.
- Strong knowledge of Azure, AWS, IBM Cloud, Docker, and Kubernetes.
- Experience with Terraform, Ansible, Puppet, The Foreman, Prometheus, Grafana, PagerDuty, and Graylog.
- Programming and scripting skills in Python, Go, Bash, or similar languages, plus expertise in CI/CD and modern deployment strategies.
- Strong analytical, problem-solving, communication, and leadership skills.
Nice to have
- Experience with large-scale distributed systems and customer-facing, high-availability production environments.
- Knowledge of networking, security, compliance best practices, incident response, and the ITIL framework.
- Understanding of core AI concepts, prompt engineering, agentic AI systems, and AI productivity tools.
Culture & Benefits
- Flexible ways of working and support for continuous learning.
- Culture focused on care, innovation, ownership, accountability, and customer success.
- Opportunity to contribute to AI-driven airline retailing and revenue optimization.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
6 дней назад
Director of Software Engineering (iGaming)
Toughbyte
14 часов назад
SRE / DevOps Engineer (GCP, Kubernetes)
5 500€
6 дней назад
Senior Observability Engineer with OpenTelemetry
4 дня назад
Senior Automation Engineer (Python/AWS)
6 дней назад
Observability Engineer (OpenTelemetry)
Kiwitaxi
14 часов назад
Backend TechLead
4 500$