11 дней назад
Lead Observability Engineer (AWS)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Lead Observability Engineer (AWS): Assessing and consolidating a large-scale observability estate across metrics, logs, and traces with an accent on OpenTelemetry, AWS-native services, cost modeling, and telemetry fidelity. Focus on measuring current-state performance, designing target architectures, leading platform migration, and rebuilding dashboards, alerts, retention, and ownership tagging.
Location: Remote; the role joins a U.S.-based Virtual Operating Center.
Company
is an embedded service provider that partners with engineering teams on complex infrastructure, cloud, and software delivery challenges.
What you will do
- Lead a two-month observability maturity assessment covering vendors, agents, collectors, query surfaces, data volumes, retention, sampling, cardinality, and operating models.
- Build defensible total cost of ownership models and compare the current estate with two to three AWS-native target states.
- Design the target architecture across OpenTelemetry, ADOT, Amazon Data Firehose, CloudWatch, Amazon Managed Service for Prometheus, Amazon Managed Grafana, S3, Athena, and OpenSearch.
- Define performance-parity tests for query performance, alert latency, and telemetry fidelity.
- Lead migration execution, including pipeline cutover, OpenTelemetry log-transform migration, dashboard and alert rebuilding, retention policies, and ownership tagging.
- Lead the embedded TechPod as a player-coach, work with customer engineering leadership, and present assessment findings, tradeoffs, and recommendations.
Requirements
- 8+ years in SRE, DevOps, observability, or platform engineering, including 4+ years owning production observability platforms and experience in a technical lead, staff, or principal role.
- Deep production experience with metrics, logs, and traces at very large scale, including millions of active series and multiple terabytes of logs per day.
- Advanced Prometheus and long-term storage experience with Thanos, Cortex, or Mimir, including PromQL, recording rules, remote write, and cardinality management.
- Production experience with AWS observability services, OpenTelemetry Collector or ADOT, Kubernetes or EKS, log pipelines, and Terraform.
- Strong scripting ability in Python, Go, or Bash and the ability to build usage-based cost models and automate migration work.
- Ability to communicate technical, cost, and compliance tradeoffs to engineering, Security, Legal, and executive stakeholders.
Nice to have
- Experience migrating from Datadog, Splunk, New Relic, or similar platforms.
- Experience automating dashboard and alert migrations with Grafana provisioning, Grafonnet, Terraform providers, or conversion tooling.
- Experience with Grafana Alloy, Beyla, eBPF instrumentation, Pyroscope, Parca, or comparable profiling tools.
- Experience with telemetry governance, SOC 2, GDPR, privacy requirements, Athena, OpenSearch, Databricks, or high-traffic B2C platforms.
- Relevant certifications such as Prometheus Certified Associate, OpenTelemetry Certified Associate, AWS professional certifications, or CKA.
Culture & Benefits
- 100% remote workplace; the company has operated remotely since its first day.
- Unlimited paid time off.
- Equity ownership.
- 401K with company contribution and sponsored healthcare.
- Training and certification programs for professional growth.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →