3 дня назад
Reliability Monitoring Engineer (Observability)
100 000 - 150 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Reliability Monitoring Engineer (Observability): Designing and operating metrics, logging, tracing, and alerting platforms for high-scale systems with an accent on signal quality, usability, and operational ROI. Focus on building telemetry pipelines, managing high-cardinality data, integrating observability with CI/CD and incident management, and applying SLOs and SRE principles.
Location: 100% remote within the United States
Salary: $100,000–$150,000 annually
Company
is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.
What you will do
- Design and operate metrics, logging, tracing, and alerting platforms across the full observability stack.
- Build and manage collection agents, telemetry pipelines, long-term storage, dashboards, and alerting workflows.
- Improve observability usability, signal quality, and operational return on investment.
- Operate high-cardinality, high-throughput metrics and log pipelines at scale.
- Integrate observability platforms with CI/CD and incident management tooling.
- Translate noisy telemetry into actionable insights for engineering and business stakeholders.
Requirements
- Must be located in the United States and have authorization to work there.
- Bachelor’s degree in Computer Science or a related field.
- At least 5 years of experience in SRE, platform engineering, or observability; the position lists 6+ years of experience.
- Hands-on experience with Prometheus, Grafana, and a major commercial observability platform such as Datadog, New Relic, or Splunk.
- Strong knowledge of OpenTelemetry, distributed tracing, structured logging, SLOs, error budgets, and SRE principles.
- Proficiency in Go, Python, or Java, plus experience with Linux internals, networking, and container platforms.
Nice to have
- Experience with Thanos, Mimir, Cortex, Loki, or Tempo at scale.
- Contributions to OpenTelemetry or other observability open-source projects.
- Familiarity with eBPF-based observability tooling.
- Experience with observability cost optimization or regulated environments requiring audit-grade logging.
Culture & Benefits
- Full-time direct W2 employment.
- 100% remote work within the United States.
- Career growth opportunities within an established technology consulting and software development organization.
- Equal employment opportunity and a workplace free from discrimination and harassment.
Hiring process
- Submit a resume for consideration by email.
- New H-1B visa petitions are not sponsored; U.S. citizens, Green Card holders, EAD holders, and H-1B transfer candidates may apply.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
2 дня назад
Staff Site Reliability Engineer (GCP/Kubernetes)
112 500 - 187 500$
5 дней назад
SRE Monitoring Platform Software Engineer (Early Career / Temporary)
21 час назад
Dynatrace Observability Consultant (DevOps)
2 дня назад
Staff Network Reliability Engineer, Cloud Operations (Kubernetes)
155 000 - 165 000$
5 дней назад
Site Reliability Expert (Observability)
100 000 - 150 000CAD
6 дней назад