5 дней назад
Principal DevOps Lead (Observability)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Principal DevOps Lead (Observability): Building and operating scalable logging, monitoring, tracing, alerting, and synthetic monitoring platforms on GCP with an accent on Kubernetes-based services, high-volume telemetry pipelines, and cloud-native observability. Focus on designing synthetic monitoring with GCP Spot VMs, establishing observability standards, and leading distributed tracing, anomaly detection, and proactive reliability initiatives.
Location: Fully remote, with the option to work flexibly from anywhere; offices in Berlin and Mannheim, Germany.
Company
provides an enterprise Conversational Cloud platform that uses conversational AI, analytics, and safety tools to support customer interactions.
What you will do
- Lead the design, implementation, operation, and continuous improvement of observability platforms for logs, metrics, traces, alerting, and synthetic monitoring.
- Own and optimize high-volume telemetry pipelines using Filebeat, Kafka, Logstash, Elastic Cloud, Prometheus, OpenTelemetry, Grafana Labs, Zabbix, and Anodot.
- Build and maintain scalable Kubernetes-based observability services with Helm, CI/CD, GCP, GKE, and Docker.
- Define observability standards, dashboards, alerting frameworks, best practices, and onboarding resources for engineering users.
- Design and develop an end-to-end synthetic monitoring platform running daily tests on GCP Spot VMs.
- Provide technical leadership and mentorship while collaborating with DevOps, SRE, Engineering, NOC, Security, and vendor teams.
Requirements
- 5+ years of experience in software engineering, DevOps, SRE, or a related discipline.
- Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent professional experience.
- Strong experience with Kubernetes, Docker, observability platforms, Grafana Labs, Zabbix, Fluentd, ELK, Kafka, and Prometheus.
- Experience with Infrastructure as Code tools such as Terraform, Ansible, or CloudFormation.
- Experience with GCP, AWS, or Azure and with designing and operating highly available, scalable, production-grade systems.
- Strong programming or scripting skills in Go, Java, JavaScript, or Python, plus CI/CD and cloud-native development experience.
Nice to have
- Experience with OpenTelemetry Collector and Grafana Agent.
- Strong problem-solving, communication, and collaboration skills with the ability to influence technical direction.
Culture & Benefits
- Remote-first model with flexible working and no defined working-hour boundaries.
- Permanent employees receive a deferred pension scheme with a 20% company contribution, ESPP, and an annual performance-based bonus.
- Internet and mobile reimbursement, professional development resources, flexible paid time off, and paid public holidays.
- 33 days of personal time off, including 28 vacation days and 5 care days.
- Volunteering days and monthly connection events at the office.
- Inclusive equal-opportunity workplace with reasonable accommodations for applicants and employees with disabilities.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
Latitude
4 дня назад
Senior Site Reliability Engineer (Kubernetes)
5 часов назад
Middle/Senior DevOps Engineer (Kubernetes)
4 дня назад
Senior DevOps Engineer (Big Data)
4 дня назад
Senior Site Reliability Engineer (Kubernetes)
6 часов назад
Observability Engineer (Terraform)
3 дня назад