Назад
7 дней назад

Site Reliability Engineer (Data and Observability Platform)

Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
China
vacancy_detail.hirify_telegram_tooltipВакансия из Telegram канала -

Мэтч & Сопровод

Покажет вашу совместимость и напишет письмо

Описание вакансии

TL;DR
Site Reliability Engineer (Data and Observability Platform): Building and operating observability, telemetry, and log pipelines for production streaming workloads with an accent on Kubernetes, AWS infrastructure, data quality, and OLAP performance. Focus on diagnosing production failures, designing actionable monitoring, recovering missed data, and maintaining reliable near-real-time analytics systems.

Site Reliability Engineer Data and Observability Platform

Company

Pulsar

Conditions

1 week ago

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will own end-to-end platform observability, build actionable dashboards and alerts, and operate production systems. You will ship changes through CI/CD and GitOps, manage Kubernetes streaming workloads and AWS infrastructure, and participate in on-call. You will also build near-real-time telemetry pipelines, implement data-quality checks, optimize OLAP tables, and write SQL for dashboards and analytics.

Requirements

  • University degree in Computer Science, Software Engineering, or a related discipline
  • At least 4–7 years of relevant work experience
  • Production observability and monitoring experience
  • CI/CD delivery experience
  • Production experience with a distributed stream-processing engine
  • Experience designing OLAP or columnar database tables and tuning queries
  • Kubernetes working knowledge
  • Strong programming skills in Python, Scala, Java, or Rust
  • Fluent analytical SQL
  • Ability to troubleshoot production systems methodically

Responsibilities

  • Own end-to-end platform observability
  • Build actionable Grafana dashboards and alerts
  • Diagnose production failures, recover missed data, and prevent recurrence
  • Ship changes through CI/CD and GitOps
  • Operate streaming workloads on Kubernetes EKS
  • Provision AWS resources using Terraform
  • Participate in on-call for the data and observability platform
  • Build and maintain near-real-time telemetry and log pipelines
  • Design data-quality checks and route bad records for investigation
  • Design and optimize OLAP tables
  • Write and tune SQL for dashboards and analytics

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →

Текст вакансии взят без изменений

Источник -