4 дня назад
Site Reliability Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer (AI) (Kubernetes, observability, SLOs): Improving the reliability, scalability, security, and operability of production infrastructure and customer-facing services running on Furiosa NPUs with an accent on reliability architecture, observability foundations, and production automation. Focus on designing SLIs and SLOs, analyzing distributed systems across software and infrastructure boundaries, and building safer rollouts, recovery mechanisms, and self-service workflows.
Location: Hybrid, Seoul, South Korea
Requirements
- Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
- Strong programming skills in one or more general-purpose languages such as Rust, Python, or Go.
- Solid understanding of operating systems, computer networks, and cloud-native or container-based environments.
- Experience improving production reliability through SLOs, observability, incident analysis, rollout safety, and error-budget-driven decisions.
- Experience designing or operating distributed systems where failures, overload, latency, and capacity limits must be explicitly managed.
- Ability to analyze technical problems and communicate clearly with engineering teams.
What you will do
- Define and evolve reliability goals through SLIs, SLOs, error budgets, and operational metrics.
- Design and build observability foundations covering metrics, logs, traces, dashboards, alerts, user impact, and failure modes.
- Analyze production systems across software, infrastructure, networking, and security boundaries, then drive architectural improvements.
- Improve rollout safety, capacity planning, load validation, graceful degradation, failure recovery, and incident learning.
- Build automation, internal tooling, and self-service workflows that reduce operational toil and improve engineering productivity.
- Operate and improve bare-metal Kubernetes clusters, cloud control planes, deployment pipelines, and API services running on Furiosa NPUs.
Culture & Benefits
- Full-time employment in a hybrid work arrangement.
- Work across production infrastructure, cloud-native systems, networking, observability, deployment, and AI accelerator services.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
4 дня назад
Senior Site Reliability Engineer (SRE)
5 дней назад
Senior Staff Service Reliability and Operational Intelligence Engineer
187 945 - 269 503$
3 дня назад
Sr. Staff Observability Engineer (GPU Cloud & Telemetry Platform)
4 дня назад
Software Engineer, Server Infrastructure (Robotics Infrastructure)
3 дня назад
Senior Site Reliability Engineer (MongoDB)
5 дней назад