Назад
Company hidden
35 минут назад

Senior Platform & Reliability Engineer (Kubernetes)

Формат работы
remote (только Germany)/hybrid/onsite
Тип работы
fulltime
Грейд
senior
Английский
c1
Страна
Germany
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Platform & Reliability Engineer (Kubernetes): Taking architectural ownership of shared infrastructure services including API gateways, persistent storage, secrets, identity, edge security, and observability with an accent on platform reliability and legacy-system modernization. Focus on designing on-call processes, building incident runbooks, rolling out tracing and SLOs/SLIs, and reducing platform-wide single points of failure.

Location: Remote in Germany; hybrid or fully on-site work is also available near offices in Berlin, Cologne, Hamburg, or Munich. Occasional travel for data center visits and team offsites is required.

Company

hirify.global provides infrastructure and hosting services supported by shared platform systems used by multiple development teams.

What you will do

  • Own the architecture and reliability of shared infrastructure services, including API gateways, ingress, persistent storage, secrets, identity, edge security, and observability.
  • Assess partially documented systems, resolve known risks, and decide which legacy components should be fixed, replaced, or retired.
  • Design and establish the on-call process, including platform and infrastructure incident runbooks.
  • Mature observability through distributed tracing, SLOs, SLIs, and meaningful dashboards.
  • Contribute system design expertise on load balancing, caching, sharding, replication, consistency, and message queues.
  • Act as a technical authority across teams while documenting and sharing critical platform knowledge.

Requirements

  • 7+ years of experience in platform, infrastructure, or SRE roles, ideally with end-to-end ownership of a private cloud or IaaS platform.
  • Production experience with Ceph and Kubernetes persistent storage such as Longhorn.
  • Experience operating API gateways and ingress, including debugging CORS and rate-limiting issues.
  • Experience with CDN and edge security, DDoS mitigation, WAF configuration, and firewall-rule design.
  • Strong knowledge of system design fundamentals and real-world architectural trade-offs.
  • Professional fluency in English is required; German is a plus.

Nice to have

  • Experience with Prometheus, Grafana, OpenTelemetry, Vault, Keycloak, NATS, and on-call processes.
  • CKA or CKS certification, Ceph training, or experience with Proxmox and OpenStack.

Culture & Benefits

  • Remote-first work with flexible hours, hybrid and office options, and equivalent technical equipment at home.
  • Workation across the EU and access to a summer office in Mallorca.
  • 30 days of vacation, additional days off on Christmas Eve and New Year's Eve, and a Volunteer Day.
  • Professional and personal development opportunities with substantial ownership and technical freedom.
  • Fitness and wellness access, corporate discounts, company events, and an international working environment.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →