12 дней назад
Senior Site Reliability Engineer (Observability)
53$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Site Reliability Engineer (Observability): Building and improving OpenTelemetry-based telemetry pipelines for Shopmonkey's production infrastructure in Google Cloud and Kubernetes with an accent on distributed tracing, metrics, logging, and service-level objectives. Focus on designing burn-rate alerts, managing observability infrastructure with Terraform and Helm, and driving incident analysis and preventative reliability improvements.
Location: Remote within Europe; reliable overlap with US Eastern Time mornings
Rate: $53 USD per hour, depending on experience
Company
is engaging an engineer for a client whose all-in-one shop management platform helps automotive businesses manage workflows, customer communications, estimates, payments, and multi-location operations.
What you will do
- Design, build, and improve OpenTelemetry pipelines for metrics, traces, and logs.
- Strengthen observability for services running on GCP and GKE, integrating Google Cloud Monitoring, Cloud Trace, Managed Service for Prometheus, and Grafana.
- Define SLIs and SLOs and implement actionable multi-window, multi-burn-rate alerts.
- Manage observability infrastructure with Terraform, Helm, and Kubernetes.
- Participate in on-call rotations, incident response, and root-cause analyses.
- Collaborate with software engineers on instrumentation and production code changes in Go or Node.js, and document operational standards.
Requirements
- 5+ years of professional experience in Site Reliability Engineering, Platform Engineering, Production Engineering, or Observability.
- Hands-on experience building production OpenTelemetry pipelines in GCP and operating Google Kubernetes Engine.
- Strong knowledge of Grafana, Prometheus, distributed tracing, metrics, logging, and telemetry architecture.
- Practical experience with SLIs, SLOs, burn-rate alerting, Terraform, Helm, and production Kubernetes infrastructure.
- Experience with on-call rotations, incident response, and thorough root-cause analysis.
- Ability to troubleshoot and submit production code changes in Go or Node.js, with strong written and verbal English communication skills.
Nice to have
- Experience personally designing observability infrastructure rather than only maintaining dashboards.
- Ability to explain tradeoffs involving sampling, cardinality, telemetry cost, alert thresholds, and signal quality.
- Experience building alerts around user impact and service reliability.
Culture & Benefits
- Hands-on, autonomous work in a fast-moving environment.
- Approximately 3.5-month contract through December 31.
- Strong potential for conversion based on performance and ongoing business needs.
- Opportunity to collaborate directly with application engineers on reliability improvements.
Hiring process
- Initial recorded screening focused on the core technical requirements.
- Technical interview with members of the engineering team.
- No take-home assessment; fast-moving process for a single hire.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
11 дней назад
Staff Site Reliability Engineer (GCP/Kubernetes)
8 дней назад
Senior SRE Engineer (Kubernetes)
5 800€
8 дней назад
Senior Site Reliability Engineer (AWS/Kubernetes)
7 000 - 12 000$
9 дней назад
Senior Site Reliability Engineer
5 000 - 9 500$
13 дней назад
Senior Site Reliability Engineer (Azure)
10 дней назад