Назад
Company hidden
3 дня назад

Sr. Staff Observability Engineer (GPU Cloud & Telemetry Platform)

Формат работы
remote (только South_korea)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
SK
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Sr. Staff Observability Engineer (GPU Cloud & Telemetry Platform) (Grafana/Prometheus): Designing and operating planet-scale telemetry pipelines for GPU-as-a-Service infrastructure with an accent on high-throughput metrics, log aggregation, distributed workloads, and GPU-specific observability. Focus on building scalable Alloy-to-Mimir and Vector-to-Loki pipelines, defining GPU SLIs/SLOs, automating reliability workflows, and correlating GPU, Kubernetes, application, and network signals for incident forensics.

Location: Seoul, South Korea

Company

hirify.global is a large global public e-commerce company building high-scale services for shopping, eating, and everyday life in South Korea.

What you will do

  • Own the end-to-end observability platform for GPU-as-a-Service infrastructure, including Grafana Alloy, Mimir, Loki, and Vector.
  • Architect low-latency, high-throughput telemetry pipelines for GPU metrics, Kubernetes and container telemetry, logs, and traces.
  • Build Grafana dashboards for GPU fleet health, tenant usage, billing insights, capacity planning, and forecasting.
  • Define GPU-specific SLIs, SLOs, error budgets, and predictive observability practices.
  • Integrate observability with CI/CD and infrastructure-as-code pipelines using Terraform and Kubernetes, including automated canary analysis and rollbacks.
  • Lead cross-layer incident forensics, design reviews, mentoring, open-source strategy, and secure multi-tenant telemetry architecture.

Requirements

  • Extensive experience in observability, SRE, or distributed infrastructure.
  • Proven experience building large-scale metrics and log telemetry pipelines.
  • Strong knowledge of Grafana Alloy or the Prometheus ecosystem, Grafana Mimir or Cortex/Thanos, Grafana Loki, and Vector or similar log pipelines.
  • Strong programming skills in Go or Python, plus experience with TSDBs and large-scale log storage.
  • Experience with Kubernetes, Linux internals, bare-metal GPU clusters, and hybrid cloud environments.
  • Experience with NVIDIA DCGM, the CUDA ecosystem, and high-performance networking such as RDMA or InfiniBand is required or preferred depending on the area.

Culture & Benefits

  • Full-time regular employment with a 12-week probation period, which may be adjusted according to business needs.
  • Entrepreneurial startup culture combined with the resources of a large public company.
  • Opportunities to drive new initiatives, innovations, and hands-on technical impact.
  • Focus on reliable, scalable infrastructure and proactive failure detection.

Hiring process

  • Application review.
  • First virtual interview followed by a second virtual interview.
  • Offer; scheduling and process details may vary by role.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →