4 дня назад
Senior DevOps Engineer (Observability)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior DevOps Engineer (Observability): Design, build, and operate a high-scale observability platform based on ClickHouse, HyperDX, OpenTelemetry, and Kubernetes with an accent on reliable telemetry pipelines, production operations, and performance optimization. Focus on building resilient log, metric, and trace processing, optimizing ClickHouse workloads, automating infrastructure, and leading complex incident investigations.
Location: Pune, India
Company
develops cybersecurity solutions and provides a production-focused environment centered on innovation and teamwork.
What you will do
- Design, build, operate, and continuously improve a high-scale observability platform using ClickHouse, HyperDX, OpenTelemetry, and Kubernetes.
- Architect and optimize ClickHouse schemas, partitioning, ordering keys, TTLs, materialized views, retention policies, storage efficiency, and queries for high-volume logs, metrics, and traces.
- Build scalable OpenTelemetry Collector pipelines with batching, queuing, retries, backpressure handling, sampling, rate limiting, enrichment, filtering, and routing.
- Operate HyperDX for log search, distributed tracing, service analysis, dashboards, telemetry correlation, troubleshooting, and root-cause analysis.
- Automate infrastructure deployment, configuration management, upgrades, application onboarding, and CI/CD workflows using Jenkins, Terraform, Ansible, Consul, Vault, and Kubernetes.
- Partner with application, platform, SRE, and operations teams on production troubleshooting, incident response, capacity planning, remediation, and observability standards.
Requirements
- 6+ years of experience in DevOps, SRE, platform, infrastructure, or observability engineering.
- Strong production experience with ClickHouse administration, architecture, schema design, performance tuning, query optimization, retention management, troubleshooting, and large-scale operations.
- Hands-on experience operating HyperDX with ClickHouse, including log search, tracing, dashboards, service analysis, and troubleshooting.
- Strong experience with OpenTelemetry Collector, Fluent Bit, Filebeat, Prometheus, Alertmanager, Grafana, Jenkins, Ansible, Terraform, Consul, Vault, and Kubernetes.
- Strong understanding of Linux, networking, microservices, REST, gRPC, distributed systems, high availability, scripting, and production operations.
- Ability to own production systems end to end, lead complex investigations, communicate during incidents, improve reliability, and mentor engineers.
Nice to have
- Java and Python development experience, including instrumentation troubleshooting and performance analysis.
- Experience building automation, internal tools, APIs, platform services, or self-service onboarding workflows.
- Experience with Kafka or other high-throughput messaging and streaming platforms.
- Experience with APM, distributed tracing, context propagation, service maps, SLOs, SLIs, alert quality, cloud platforms, and large-scale Kubernetes environments.
Culture & Benefits
- Hands-on ownership of critical observability services and production outcomes.
- Automation-first engineering practices focused on repeatability, reliability, scalability, performance, cost, and operational safety.
- Technical leadership through design reviews, documentation, runbooks, incident reviews, standards development, and mentoring.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
20 часов назад
Dynatrace Observability Consultant (DevOps)
17 часов назад
Senior Observability Engineer (AI)
2 дня назад
Senior DevOps Engineer (Kubernetes)
4 дня назад
Observability Engineer (Cloud)
100 000 - 160 000$
5 дней назад
Senior DevOps / SRE Engineer
5 500$
4 дня назад
Site Observability Engineer
100 000 - 150 000$