1 день назад
Observability Engineer (Cloud)
100 000 - 160 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Observability Engineer (Cloud): Building and operating enterprise observability platforms for metrics, logs, traces, events, and synthetic monitoring with an accent on OpenTelemetry, high-scale telemetry pipelines, and SLO-driven operations. Focus on architecting Prometheus, Grafana, Loki, Tempo, Thanos, Mimir, and Datadog deployments, reducing alert noise, optimizing storage costs, and integrating observability with incident response and delivery workflows.
Location: 100% remote within the United States
Salary: $100,000–$160,000 annually
Company
is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.
What you will do
- Design and operate enterprise-grade observability platforms covering metrics, logs, traces, events, and synthetic monitoring.
- Architect highly available Prometheus, Thanos, Mimir, Grafana, Loki, Tempo, OpenTelemetry, and Datadog deployments.
- Define service instrumentation standards, SLOs, SLIs, error budgets, dashboards, and actionable alerting strategies.
- Operate large-scale time-series and log storage platforms while balancing retention, query performance, and cost.
- Build self-service tooling, libraries, templates, documentation, and runbooks to support observability adoption.
- Partner with SRE and platform teams on incident readiness, CI/CD integration, canary analysis, progressive delivery, and post-incident improvements.
Requirements
- Bachelor’s degree in Computer Science or a related field.
- 10 or more years of experience in SRE, platform engineering, or observability roles.
- Hands-on experience with Prometheus, Grafana, and at least one commercial observability platform such as Datadog, New Relic, or Splunk.
- Strong knowledge of OpenTelemetry, distributed tracing, structured logging, SLOs, error budgets, and SRE principles.
- Proficiency in Go, Python, or Java, plus experience with high-cardinality metrics and high-throughput log pipelines.
- Experience with CI/CD and incident management tooling, Linux internals, networking, and container platforms.
Nice to have
- Experience with Thanos, Mimir, Cortex, Loki, or Tempo at scale.
- Contributions to OpenTelemetry or other observability open-source projects.
- Familiarity with eBPF-based observability tooling and observability cost optimization.
- Experience with regulated environments and audit-grade logging requirements.
Culture & Benefits
- Full-time direct W-2 employment.
- 100% remote work within the United States.
- Career growth opportunities within an established consulting and software development organization.
- Equal employment opportunity and a workplace free from discrimination and harassment.
Hiring process
- Submit a resume for consideration.
- New H-1B visa petitions are not sponsored; U.S. citizens, Green Card holders, EAD holders, and H-1B transfer candidates are encouraged to apply.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
3 дня назад
Senior DevOps Engineer (Kubernetes)
140 000 - 192 500CAD
Reddit
2 дня назад
Staff Software Engineer (Observability)
217 000 - 303 900$
2 дня назад
Site Reliability Engineer, USG (Aerospace)
6 дней назад
Senior AWS DevOps Engineer
91 200 - 120 000$
4 дня назад
Senior Software Engineer – DevOps/Observability Platform
55 500 - 93 600€
2 дня назад
Site Reliability Expert (Observability)
100 000 - 150 000CAD