7 часов назад
Technical Incident Commander
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Technical Incident Commander (Live Operations): Managing critical S1–S3 incidents and restoring services across DAZN’s global technology infrastructure with an accent on incident coordination, operational readiness, and customer-impact analysis. Focus on leading high-pressure war rooms, improving observability and runbooks, and reducing MTTR across a 24/7 live sports streaming operation.
Location: Leeds, UK; hybrid workplace
Company
operates a global sports streaming platform delivering live and on-demand content to millions of fans.
What you will do
- Lead the technical resolution and service restoration of critical S1–S3 incidents.
- Coordinate engineering, operations, support, subject-matter experts, and third-party vendors through war rooms and structured communications.
- Own the Major Incident Management process and ensure SLAs and response procedures are followed.
- Communicate incident updates to business stakeholders while enabling engineering teams to focus on restoration.
- Operationalize new features by defining monitoring requirements, producing runbooks, and assessing delivery risks.
- Identify infrastructure, architecture, tooling, and process improvements that increase uptime and reduce MTTR.
Requirements
- Proven experience managing and resolving critical S1–S3 incidents.
- Strong understanding of ITIL incident, problem, and change management practices.
- Experience leading cross-functional and vendor teams during high-pressure incidents.
- Background in 24/7/365 operations and experience analyzing incident data, including support tickets, social signals, error codes, and playback failures.
- Hands-on knowledge of Halo, JIRA, APIs, cloud platforms, ECS, Lambda, and cloud-native databases.
- Experience transitioning complex projects into live environments with monitoring, runbooks, and operational frameworks.
Nice to have
- Experience with observability tools such as New Relic, Coralogix, or Conviva.
- Familiarity with customer-impact KPIs including rebuffering rates, playback failures, and capacity alerts.
Culture & Benefits
- Permanent, full-time employment in a hybrid workplace.
- Participation in a 1-in-4 on-call rotation supporting major live events.
- 25 days of annual leave, increasing by 3 days after three years.
- Private medical insurance, life assurance, and pension contributions of up to 5%.
- Enhanced parental leave, an electric vehicle benefit option, and learning and development resources.
- Access to , internal speaker series, events, and flexible working options.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
4 дня назад
Engineering Manager, Edge SRE
3 дня назад
Senior Site Reliability Engineer (Kubernetes)
Nscale
3 дня назад
Centre of Excellence Senior Engineer (AI Infrastructure)
2 дня назад
Director, Site Reliability Engineering
87 000 - 163 000GBP
2 дня назад
Sr. Manager, Site Reliability
59 550 - 110 594GBP
6 дней назад