3 дня назад
Sr. Staff Observability Engineer (GPU Cloud & Telemetry Platform)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Sr. Staff Observability Engineer (GPU Cloud & Telemetry Platform) (Grafana/Prometheus): Designing and operating planet-scale telemetry pipelines for GPU-as-a-Service infrastructure with an accent on high-throughput metrics, log aggregation, distributed workloads, and GPU-specific observability. Focus on building scalable Alloy-to-Mimir and Vector-to-Loki pipelines, defining GPU SLIs/SLOs, automating reliability workflows, and correlating GPU, Kubernetes, application, and network signals for incident forensics.
Location: Seoul, South Korea
Company
is a large global public e-commerce company building high-scale services for shopping, eating, and everyday life in South Korea.
What you will do
- Own the end-to-end observability platform for GPU-as-a-Service infrastructure, including Grafana Alloy, Mimir, Loki, and Vector.
- Architect low-latency, high-throughput telemetry pipelines for GPU metrics, Kubernetes and container telemetry, logs, and traces.
- Build Grafana dashboards for GPU fleet health, tenant usage, billing insights, capacity planning, and forecasting.
- Define GPU-specific SLIs, SLOs, error budgets, and predictive observability practices.
- Integrate observability with CI/CD and infrastructure-as-code pipelines using Terraform and Kubernetes, including automated canary analysis and rollbacks.
- Lead cross-layer incident forensics, design reviews, mentoring, open-source strategy, and secure multi-tenant telemetry architecture.
Requirements
- Extensive experience in observability, SRE, or distributed infrastructure.
- Proven experience building large-scale metrics and log telemetry pipelines.
- Strong knowledge of Grafana Alloy or the Prometheus ecosystem, Grafana Mimir or Cortex/Thanos, Grafana Loki, and Vector or similar log pipelines.
- Strong programming skills in Go or Python, plus experience with TSDBs and large-scale log storage.
- Experience with Kubernetes, Linux internals, bare-metal GPU clusters, and hybrid cloud environments.
- Experience with NVIDIA DCGM, the CUDA ecosystem, and high-performance networking such as RDMA or InfiniBand is required or preferred depending on the area.
Culture & Benefits
- Full-time regular employment with a 12-week probation period, which may be adjusted according to business needs.
- Entrepreneurial startup culture combined with the resources of a large public company.
- Opportunities to drive new initiatives, innovations, and hands-on technical impact.
- Focus on reliable, scalable infrastructure and proactive failure detection.
Hiring process
- Application review.
- First virtual interview followed by a second virtual interview.
- Offer; scheduling and process details may vary by role.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
2 дня назад
Senior SRE (Security/Telemetry)
4 дня назад
Site Reliability Engineer (AI)
3 дня назад
Site Reliability Engineer (AI)
4 дня назад
Site Reliability Engineer Staff (Cloud Infrastructure)
2 дня назад
Senior SRE (Kubernetes/GCP)
4 дня назад
Senior Production Operations Engineer (AWS/Kubernetes)
172 100 - 258 100$