Назад
Company hidden
11 дней назад

Lead Observability Engineer (AWS)

Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
lead
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Lead Observability Engineer (AWS): Assessing and consolidating a large-scale observability estate across metrics, logs, and traces with an accent on OpenTelemetry, AWS-native services, cost modeling, and telemetry fidelity. Focus on measuring current-state performance, designing target architectures, leading platform migration, and rebuilding dashboards, alerts, retention, and ownership tagging.

Location: Remote; the role joins a U.S.-based Virtual Operating Center.

Company

hirify.global is an embedded service provider that partners with engineering teams on complex infrastructure, cloud, and software delivery challenges.

What you will do

  • Lead a two-month observability maturity assessment covering vendors, agents, collectors, query surfaces, data volumes, retention, sampling, cardinality, and operating models.
  • Build defensible total cost of ownership models and compare the current estate with two to three AWS-native target states.
  • Design the target architecture across OpenTelemetry, ADOT, Amazon Data Firehose, CloudWatch, Amazon Managed Service for Prometheus, Amazon Managed Grafana, S3, Athena, and OpenSearch.
  • Define performance-parity tests for query performance, alert latency, and telemetry fidelity.
  • Lead migration execution, including pipeline cutover, OpenTelemetry log-transform migration, dashboard and alert rebuilding, retention policies, and ownership tagging.
  • Lead the embedded TechPod as a player-coach, work with customer engineering leadership, and present assessment findings, tradeoffs, and recommendations.

Requirements

  • 8+ years in SRE, DevOps, observability, or platform engineering, including 4+ years owning production observability platforms and experience in a technical lead, staff, or principal role.
  • Deep production experience with metrics, logs, and traces at very large scale, including millions of active series and multiple terabytes of logs per day.
  • Advanced Prometheus and long-term storage experience with Thanos, Cortex, or Mimir, including PromQL, recording rules, remote write, and cardinality management.
  • Production experience with AWS observability services, OpenTelemetry Collector or ADOT, Kubernetes or EKS, log pipelines, and Terraform.
  • Strong scripting ability in Python, Go, or Bash and the ability to build usage-based cost models and automate migration work.
  • Ability to communicate technical, cost, and compliance tradeoffs to engineering, Security, Legal, and executive stakeholders.

Nice to have

  • Experience migrating from Datadog, Splunk, New Relic, or similar platforms.
  • Experience automating dashboard and alert migrations with Grafana provisioning, Grafonnet, Terraform providers, or conversion tooling.
  • Experience with Grafana Alloy, Beyla, eBPF instrumentation, Pyroscope, Parca, or comparable profiling tools.
  • Experience with telemetry governance, SOC 2, GDPR, privacy requirements, Athena, OpenSearch, Databricks, or high-traffic B2C platforms.
  • Relevant certifications such as Prometheus Certified Associate, OpenTelemetry Certified Associate, AWS professional certifications, or CKA.

Culture & Benefits

  • 100% remote workplace; the company has operated remotely since its first day.
  • Unlimited paid time off.
  • Equity ownership.
  • 401K with company contribution and sponsored healthcare.
  • Training and certification programs for professional growth.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →