Назад
Company hidden
6 дней назад

Site Reliability Engineer (Kafka/Kubernetes)

Формат работы
hybrid
Тип работы
fulltime
Грейд
middle
Английский
b2
Страна
Germany
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Site Reliability Engineer (Kafka/Kubernetes): Operating trivago’s high-throughput data backbone across Kafka, Flink, Cassandra, Redis, Google Cloud Platform, and on-premises infrastructure with an accent on reliability engineering, observability, and automation. Focus on debugging distributed data-intensive systems, improving incident response, and building AI-assisted operational tooling.

Location: Düsseldorf, Germany; hybrid work with up to 2 work-from-home days weekly

Company

hirify.global operates a hotel metasearch engine that enables travelers to compare prices from hundreds of booking sites.

What you will do

  • Operate 8 production Kafka clusters processing more than 900K messages per second, along with managed Flink, Cassandra, and Redis services.
  • Run infrastructure across Google Cloud Platform, including GKE and Compute Engine, and an on-premises datacenter using Rancher.
  • Participate in the on-call rotation and support incident recovery across data services.
  • Build monitoring and observability for applications and infrastructure.
  • Document operational procedures, create runbooks, and automate repetitive work with AI-assisted tooling where useful.
  • Debug issues across multiple services and contribute to engineering issues and OKRs.

Requirements

  • Hands-on experience operating Kubernetes in production.
  • Practical experience with Kafka.
  • Experience running infrastructure in both cloud and on-premises environments; Google Cloud Platform, GKE, Compute Engine, and Rancher are relevant technologies.
  • Strong analytical and troubleshooting skills in distributed, data-intensive systems.
  • Ability to prioritize work independently, including during incident pressure.
  • Curiosity about AI tools and automation for incident triage, runbook automation, and daily operations.

Nice to have

  • Experience with Flink, Cassandra, Redis, or MySQL.
  • Interest in contributing to technical documentation and engineering blog content.

Culture & Benefits

  • Small autonomous teams with direct access to production and ownership of technical decisions.
  • Continuous deployment, short feedback loops, and dedicated experimentation time.
  • Technical learning opportunities including meetups, guilds, architecture reviews, certifications, conferences, workshops, online courses, and a campus library.
  • Visa support, a relocation package, an interest-free newcomer loan, and free language classes.
  • Self-determined vacation with a minimum of 25 days, flexible working hours, and the option to work remotely from Germany or selected countries abroad for up to 20 days per year.
  • Daily canteen budget, snacks and drinks, an on-site gym, sports classes, Urban Sports Club membership, ergonomic workspaces, childcare support, and an on-campus kids room.

Hiring process

  • Asynchronous video interview.
  • Case study followed by a technical interview and case study presentation.
  • Profile Dynamics Assessment and final interview with key stakeholders.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →