14 дней назад
Resilience Engineer
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Resilience Engineer (SRE/Observability): Developing and governing resilience strategies for highly available IoT systems with an accent on fault tolerance, observability, chaos engineering, and disaster recovery. Focus on designing failover and degraded-mode scenarios, automating fault detection, and applying AI-assisted analysis to improve reliability.
Location: Lisboa, Portugal. Flexible hybrid work model with 8–10 in-office days per month.
Company
is an international telecommunications company providing connectivity and technology services to millions of customers.
What you will do
- Develop and govern resilience strategies across system architecture, deployment, monitoring, and incident response.
- Define and track stability KPIs, including MTTD, MTTR, SLOs, SLIs, and error budgets.
- Design fault-injection tests, chaos engineering practices, and scenario-based simulations.
- Collaborate with product, infrastructure, architecture, and development teams to implement redundancy, failover, and graceful degradation.
- Drive automation and observability improvements for faster fault detection, reduced noise, and predictive failure mitigation.
- Maintain Business Continuity and Disaster Recovery plans for resilient and recoverable IoT systems.
Requirements
- BSc degree in Software Engineering, Computer Science, or a related discipline, or equivalent professional experience.
- Strong expertise in Site Reliability Engineering, distributed systems, observability, systemic reliability improvement, and dependency management.
- Experience analysing performance, failure modes, latency, throughput, errors, saturation, service degradation, failover, and recovery scenarios.
- Knowledge of Grafana, Prometheus, OpenTelemetry, Loki, Splunk, synthetic monitoring, and service-level performance monitoring.
- Experience applying AI-assisted observability and automation to anomaly detection, incident analysis, event correlation, root-cause investigation, and predictive reliability with human oversight.
- Fluent written and spoken English required. Strong analytical, communication, stakeholder-management, influencing, ownership, and prioritisation skills.
Nice to have
- Experience with resilience testing, failure injection, or Chaos Engineering.
- Understanding of Business Continuity and Disaster Recovery standards such as ISO 22301.
- Experience with Real User Monitoring.
Culture & Benefits
- Flexible hybrid work model managed by team leaders.
- Mobile phone, communication plan, data card, and discounts on products and services.
- Recognition programs for innovative, creative, and high-potential contributions.
- Well-being support including nutrition and psychological consultations, webinars, workshops, and discounts.
- Access to Communities of Practice and a digital training platform with professional learning content.
- Local and international internal mobility opportunities across departments and roles.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →