7 дней назад
Senior Director of Production Engineering (Observability & Telemetry Platforms)
231 000 - 330 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Director of Production Engineering (Observability & Telemetry Platforms) (Observability/Telemetry): Defining and scaling global observability platforms and high-throughput telemetry pipelines for a hyperscale cloud security platform with an accent on OpenTelemetry modernization, petabyte-scale streaming, and platform reliability. Focus on designing fault-tolerant architectures, standardizing SLOs and error budgets, automating incident response toil, and leading distributed engineering organizations.
Location: San Jose, California, USA; hybrid with at least 3 days per week onsite at headquarters
Base salary: $231,000–$330,000 USD annually, excluding commission, bonus, equity, and benefits.
Company
provides a cloud security platform based on Zero Trust Exchange and SASE technologies, protecting customers from cyberattacks and data loss.
What you will do
- Define the long-term roadmap and architectural strategy for the global Observability Platform.
- Modernize monitoring stacks using open standards such as OpenTelemetry and optimize total cost of ownership.
- Architect and scale fault-tolerant, low-latency telemetry ingestion and processing pipelines handling petabyte-scale streaming data.
- Optimize platform performance, resource utilization, durability, and cost efficiency as telemetry volumes grow.
- Establish standardized SLOs, error budgets, high-fidelity alerting, and streamlined incident escalation workflows with SRE and incident management teams.
- Lead and develop a distributed engineering organization while partnering with Product Management, Information Security, and Customer Support.
Requirements
- 7+ years of progressive engineering leadership experience managing managers, technical leads, and distributed organizations of 30+ engineers.
- Experience in hyperscale SaaS, cloud provider, or high-transaction distributed systems environments.
- Expertise in observability and telemetry systems, including distributed tracing, high-cardinality metrics, distributed log analytics, and telemetry routing.
- Experience with OpenTelemetry, Kafka, Flink, Vector, VictoriaMetrics, Grafana, Prometheus, ClickHouse, or OpenSearch.
- Strong foundation in cloud-native architecture, Kubernetes, microservices, SRE methodologies, and incident command structures.
- Ability to work onsite at headquarters in San Jose at least 3 days per week.
Nice to have
- Experience with generative AI, AI/ML models, or intelligent agents for production engineering and predictive observability.
- Experience with compliance, data retention, PII masking, and data governance in telemetry environments.
- Open-source contributions to observability projects such as OpenTelemetry or Prometheus.
- Experience with chaos engineering or automated disaster recovery testing at scale.
Culture & Benefits
- Inclusive environment focused on ownership, collaboration, transparency, psychological safety, and continuous learning.
- Health plans, vacation and sick time, parental leave, retirement options, and education reimbursement.
- In-office perks and support for employees and their families.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
8 дней назад
Senior Engineering Manager, Observability
220 000 - 303 000$
7 дней назад
Site Observability Engineer
100 000 - 150 000$
10 дней назад
Telemetry Data Infrastructure DevOps Software Engineer III (AI/ML)
164 652 - 230 512$
9 дней назад
Software Engineering Manager – Remote Services (AI)
153 300 - 260 600$
10 дней назад
Senior Infrastructure & DevOps Engineer
195 000 - 235 000$
10 дней назад
Manager, Solutions Engineering (Observability)
234 000 - 278 000$