4 дня назад
Site Observability Engineer
100 000 - 150 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Observability Engineer (Prometheus/Grafana/OpenTelemetry): Designing and operating enterprise observability platforms for metrics, logs, traces, events, and synthetic monitoring with an accent on signal quality, high availability, and operational ROI. Focus on building scalable telemetry pipelines, defining SLOs and error budgets, optimizing storage and alerting, and integrating observability with incident response and progressive delivery.
Location: 100% remote within the United States
Salary: $100,000–$150,000 annually
Company
Technology consulting and software development organization delivering cloud, AI, data, and enterprise solutions across the United States.
What you will do
- Design and operate enterprise observability platforms for metrics, logs, traces, events, and synthetic monitoring.
- Architect highly available Prometheus, Thanos, Mimir, Grafana, Loki, Tempo, OpenTelemetry, and Datadog deployments.
- Define service instrumentation standards, SLOs, SLIs, error budgets, dashboards, and actionable alerting strategies.
- Operate large-scale time-series and log storage while balancing retention, query performance, and cost.
- Build self-service tooling, libraries, templates, and documentation that support observability adoption across product teams.
- Partner with SRE and platform teams on incident readiness, CI/CD integration, canary analysis, and progressive delivery.
Requirements
- Must be located in the United States.
- Bachelor’s degree in Computer Science or a related field.
- At least five years of experience in SRE, platform engineering, or observability roles; the position listing specifies six or more years of experience.
- Hands-on experience with Prometheus, Grafana, and a major commercial observability platform such as Datadog, New Relic, or Splunk.
- Strong knowledge of OpenTelemetry, distributed tracing, structured logging, SLOs, error budgets, and SRE principles.
- Proficiency in Go, Python, or Java, plus experience with Linux, networking, containers, high-throughput telemetry pipelines, CI/CD, and incident management tools.
Nice to have
- Experience with Thanos, Mimir, Cortex, Loki, or Tempo at scale.
- Contributions to OpenTelemetry or other observability open-source projects.
- Familiarity with eBPF-based observability tooling.
- Experience with observability cost optimization or regulated environments requiring audit-grade logging.
Culture & Benefits
- Full-time direct W-2 employment.
- Career growth opportunities within an established organization.
- Collaboration with engineering, SRE, platform, and product teams.
- Applicants must be U.S. citizens, Green Card holders, EAD holders, or H-1B transfer candidates; new H-1B visa petitions cannot be sponsored.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
2 дня назад
Staff Network Reliability Engineer, Cloud Operations (Kubernetes)
155 000 - 165 000$
2 дня назад
Staff Site Reliability Engineer (GCP/Kubernetes)
112 500 - 187 500$
6 дней назад
Senior Site Reliability Engineer (Observability)
22 часа назад
Site Reliability Engineer (SRE)
5 дней назад
Principal Production Engineer
160 200 - 425 000$
5 дней назад
Site Reliability Expert (Observability)
100 000 - 150 000CAD