Назад
Company hidden
4 дня назад

Incident Management Reliability Engineer (AIOps)

49 600 - 74 400€
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
c1
Страна
Spain
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Incident Management Reliability Engineer (AIOps): Managing major incidents and improving the reliability, observability, and fault tolerance of critical IT services for a global biopharmaceutical company with an accent on incident command, technical troubleshooting, and data-driven recovery. Focus on defining SLOs, automating recovery, analyzing failure modes, and eliminating recurring root causes through AIOps and SRE practices.

Location: Barcelona, Spain; hybrid

Salary: €49,600–€74,400 per year

Company

hirify.global is a global biopharmaceutical company developing medicines and vaccines to improve human health.

What you will do

  • Lead the end-to-end management of P1/P2 major incidents and coordinate command-center response during critical outages.
  • Provide technical leadership during high-severity incidents, bringing together technical and business stakeholders to accelerate diagnosis and recovery.
  • Maintain incident documentation, lead post-incident reviews, and ensure action items prevent recurrence.
  • Improve service reliability, observability, fault tolerance, monitoring, alerting, and automated recovery with service owners and platform teams.
  • Analyze incident trends, contribute to capacity planning and failure mode analysis, and define SLOs, SLIs, and SLAs.
  • Automate operational tasks and evolve major incident processes using scripts, runbooks, AIOps, and SRE practices.

Requirements

  • 8+ years of experience in incident management, reliability engineering, or a related IT operations discipline.
  • Strong foundation in networking and infrastructure as code.
  • Experience with cloud technologies, virtualization, containerization, automation, databases, or middleware is beneficial.
  • Preferred certifications include ITIL v4, SRE Foundation or Practitioner, AWS, Azure, GCP, or Incident Command System training.
  • Fluent English, written and verbal, is required.
  • Excellent communication, collaboration, and decision-making skills under pressure.

Culture & Benefits

  • Work on technology supporting life-changing medicines and vaccines.
  • Access learning opportunities and professional certification support.
  • Collaborative, service-focused environment centered on resilience and continuous improvement.
  • Opportunity to work with AIOps, automation, and modern SRE practices in an enterprise environment.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →