Назад
Company hidden
7 дней назад

Senior Director of Production Engineering (Observability & Telemetry Platforms)

231 000 - 330 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior/director
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Director of Production Engineering (Observability & Telemetry Platforms) (Observability/Telemetry): Defining and scaling global observability platforms and high-throughput telemetry pipelines for a hyperscale cloud security platform with an accent on OpenTelemetry modernization, petabyte-scale streaming, and platform reliability. Focus on designing fault-tolerant architectures, standardizing SLOs and error budgets, automating incident response toil, and leading distributed engineering organizations.

Location: San Jose, California, USA; hybrid with at least 3 days per week onsite at headquarters

Base salary: $231,000–$330,000 USD annually, excluding commission, bonus, equity, and benefits.

Company

hirify.global provides a cloud security platform based on Zero Trust Exchange and SASE technologies, protecting customers from cyberattacks and data loss.

What you will do

  • Define the long-term roadmap and architectural strategy for the global Observability Platform.
  • Modernize monitoring stacks using open standards such as OpenTelemetry and optimize total cost of ownership.
  • Architect and scale fault-tolerant, low-latency telemetry ingestion and processing pipelines handling petabyte-scale streaming data.
  • Optimize platform performance, resource utilization, durability, and cost efficiency as telemetry volumes grow.
  • Establish standardized SLOs, error budgets, high-fidelity alerting, and streamlined incident escalation workflows with SRE and incident management teams.
  • Lead and develop a distributed engineering organization while partnering with Product Management, Information Security, and Customer Support.

Requirements

  • 7+ years of progressive engineering leadership experience managing managers, technical leads, and distributed organizations of 30+ engineers.
  • Experience in hyperscale SaaS, cloud provider, or high-transaction distributed systems environments.
  • Expertise in observability and telemetry systems, including distributed tracing, high-cardinality metrics, distributed log analytics, and telemetry routing.
  • Experience with OpenTelemetry, Kafka, Flink, Vector, VictoriaMetrics, Grafana, Prometheus, ClickHouse, or OpenSearch.
  • Strong foundation in cloud-native architecture, Kubernetes, microservices, SRE methodologies, and incident command structures.
  • Ability to work onsite at headquarters in San Jose at least 3 days per week.

Nice to have

  • Experience with generative AI, AI/ML models, or intelligent agents for production engineering and predictive observability.
  • Experience with compliance, data retention, PII masking, and data governance in telemetry environments.
  • Open-source contributions to observability projects such as OpenTelemetry or Prometheus.
  • Experience with chaos engineering or automated disaster recovery testing at scale.

Culture & Benefits

  • Inclusive environment focused on ownership, collaboration, transparency, psychological safety, and continuous learning.
  • Health plans, vacation and sick time, parental leave, retirement options, and education reimbursement.
  • In-office perks and support for employees and their families.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →