Назад
Company hidden
14 дней назад

Resilience Engineer

Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
c1
Страна
Portugal
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Resilience Engineer (SRE/Observability): Developing and governing resilience strategies for highly available IoT systems with an accent on fault tolerance, observability, chaos engineering, and disaster recovery. Focus on designing failover and degraded-mode scenarios, automating fault detection, and applying AI-assisted analysis to improve reliability.

Location: Lisboa, Portugal. Flexible hybrid work model with 8–10 in-office days per month.

Company

hirify.global is an international telecommunications company providing connectivity and technology services to millions of customers.

What you will do

  • Develop and govern resilience strategies across system architecture, deployment, monitoring, and incident response.
  • Define and track stability KPIs, including MTTD, MTTR, SLOs, SLIs, and error budgets.
  • Design fault-injection tests, chaos engineering practices, and scenario-based simulations.
  • Collaborate with product, infrastructure, architecture, and development teams to implement redundancy, failover, and graceful degradation.
  • Drive automation and observability improvements for faster fault detection, reduced noise, and predictive failure mitigation.
  • Maintain Business Continuity and Disaster Recovery plans for resilient and recoverable IoT systems.

Requirements

  • BSc degree in Software Engineering, Computer Science, or a related discipline, or equivalent professional experience.
  • Strong expertise in Site Reliability Engineering, distributed systems, observability, systemic reliability improvement, and dependency management.
  • Experience analysing performance, failure modes, latency, throughput, errors, saturation, service degradation, failover, and recovery scenarios.
  • Knowledge of Grafana, Prometheus, OpenTelemetry, Loki, Splunk, synthetic monitoring, and service-level performance monitoring.
  • Experience applying AI-assisted observability and automation to anomaly detection, incident analysis, event correlation, root-cause investigation, and predictive reliability with human oversight.
  • Fluent written and spoken English required. Strong analytical, communication, stakeholder-management, influencing, ownership, and prioritisation skills.

Nice to have

  • Experience with resilience testing, failure injection, or Chaos Engineering.
  • Understanding of Business Continuity and Disaster Recovery standards such as ISO 22301.
  • Experience with Real User Monitoring.

Culture & Benefits

  • Flexible hybrid work model managed by team leaders.
  • Mobile phone, communication plan, data card, and discounts on hirify.global products and services.
  • Recognition programs for innovative, creative, and high-potential contributions.
  • Well-being support including nutrition and psychological consultations, webinars, workshops, and discounts.
  • Access to Communities of Practice and a digital training platform with professional learning content.
  • Local and international internal mobility opportunities across departments and roles.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →