4 дня назад
Site Reliability Engineer
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer (Cloud/DevOps): Improving and scaling the reliability, security, performance, and efficiency of internal and hosted production products with an accent on observability, incident management, infrastructure as code, and cloud-native architecture. Focus on defining SLIs and SLOs, building reusable platform components, automating operations with an AI-first approach, and designing resilient systems at scale.
Location: Belgium
Company
is an independent European digital product studio delivering digital products, strategies, and software for enterprises and governments.
What you will do
- Take shared ownership of internal and hosted production products, improving reliability, security, performance, and cost efficiency.
- Participate in the full incident lifecycle, including detection, response, mitigation, resolution, and blameless post-mortems, as part of a sustainable on-call rotation.
- Define and implement SLIs, SLOs, and error budgets, and provide technical input for maintenance, SLA presales, proposals, and contracts.
- Develop observability capabilities with a strong focus on OpenTelemetry and build reusable platform components such as GitLab components and Terraform modules.
- Work hands-on with CI/CD, infrastructure as code, cloud-native architecture, networking, resilience, performance, and scalability.
- Embed security through secrets management, SAST/DAST scanning, penetration testing, vulnerability remediation, and automation using an AI-first operations approach.
Requirements
- Hands-on experience as a DevOps or SRE engineer running production products and taking ownership of their reliability.
- Experience with incident management and some exposure to SLIs and SLOs.
- Hands-on foundation across cloud-native platforms such as GCP or AWS, CI/CD, GitLab, Terraform, networking, observability, and resilience.
- Security-minded approach covering SAST/DAST, penetration testing, secrets management, and vulnerability management.
- Interest in platform engineering, self-service capabilities, reusable components, and inner sourcing.
- Collaborative communication with development teams, client partners, and customer operations teams.
Nice to have
- Experience with incident.io and OpenTelemetry.
- Experience with maintenance or SLA contracts and technical presales.
Culture & Benefits
- Permanent contract with responsibility and the opportunity to make an impact from day one.
- Hospitalisation and group insurance, meal vouchers, phone subscription, and other benefits.
- Mobility budget with flexibility in choosing transportation.
- Team events throughout the year.
- Dedicated learning and development budget after six months.
- Multidisciplinary, autonomous teams with opportunities to learn, take ownership, and grow.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →