Назад
Company hidden
4 дня назад

Site Reliability Engineer (AI)

Формат работы
hybrid
Тип работы
fulltime
Английский
b2
Страна
SK
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Site Reliability Engineer (AI) (Kubernetes, observability, SLOs): Improving the reliability, scalability, security, and operability of production infrastructure and customer-facing services running on Furiosa NPUs with an accent on reliability architecture, observability foundations, and production automation. Focus on designing SLIs and SLOs, analyzing distributed systems across software and infrastructure boundaries, and building safer rollouts, recovery mechanisms, and self-service workflows.

Location: Hybrid, Seoul, South Korea

Requirements

  • Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
  • Strong programming skills in one or more general-purpose languages such as Rust, Python, or Go.
  • Solid understanding of operating systems, computer networks, and cloud-native or container-based environments.
  • Experience improving production reliability through SLOs, observability, incident analysis, rollout safety, and error-budget-driven decisions.
  • Experience designing or operating distributed systems where failures, overload, latency, and capacity limits must be explicitly managed.
  • Ability to analyze technical problems and communicate clearly with engineering teams.

What you will do

  • Define and evolve reliability goals through SLIs, SLOs, error budgets, and operational metrics.
  • Design and build observability foundations covering metrics, logs, traces, dashboards, alerts, user impact, and failure modes.
  • Analyze production systems across software, infrastructure, networking, and security boundaries, then drive architectural improvements.
  • Improve rollout safety, capacity planning, load validation, graceful degradation, failure recovery, and incident learning.
  • Build automation, internal tooling, and self-service workflows that reduce operational toil and improve engineering productivity.
  • Operate and improve bare-metal Kubernetes clusters, cloud control planes, deployment pipelines, and API services running on Furiosa NPUs.

Culture & Benefits

  • Full-time employment in a hybrid work arrangement.
  • Work across production infrastructure, cloud-native systems, networking, observability, deployment, and AI accelerator services.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →