Назад
Company hidden
4 дня назад

Senior Site Reliability Engineer (SRE)

Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
SK
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Site Reliability Engineer (SRE) (Cloud Infrastructure): Building and scaling reliability practices, observability systems, and automation for distributed cloud-native services with an accent on SLOs, incident response, and production performance. Focus on designing actionable monitoring, reducing operational toil, leading complex incidents, and establishing reusable reliability standards across engineering.

Location: Seoul, South Korea; must be a South Korean citizen or currently reside in South Korea with a valid work visa

Company

hirify.global develops an AI-powered, cloud-native contact center platform focused on multimodal customer experiences, automation, scalability, and security.

What you will do

  • Lead improvements to reliability, scalability, and performance across critical distributed services.
  • Define and implement SLIs, SLOs, and error budgets to guide engineering priorities.
  • Design observability systems covering metrics, logging, tracing, and alerting.
  • Lead complex incident response, serve as incident commander when needed, and drive systemic postmortem actions.
  • Reduce operational toil through automation, tooling, improved workflows, and reusable operational systems.
  • Partner with product and platform teams on architecture, production readiness, failure recovery, and reliability practices while mentoring engineers.

Requirements

  • South Korean citizenship or current residence in South Korea with a valid work visa
  • 6–10+ years of experience in SRE, infrastructure, or backend systems engineering.
  • Experience owning reliability outcomes for complex distributed systems.
  • Strong experience with cloud infrastructure such as AWS, GCP, or Azure and production-scale systems.
  • Deep understanding of observability, incident management, and system performance.
  • Proficiency in at least one programming language, such as Go, Python, or Java, focused on automation and tooling.

Nice to have

  • Experience building or scaling SRE practices, including SLOs, incident frameworks, and on-call models.
  • Kubernetes or container orchestration experience.
  • Infrastructure as Code experience, including Terraform.
  • Experience with high-growth systems, performance engineering, or capacity planning.

Culture & Benefits

  • Early opportunity to shape the company’s SRE function, reliability standards, and engineering practices.
  • Collaborative and inclusive environment valuing creative solutions and strong working relationships.
  • Medical, dental, and vision benefits.
  • 401(k) plan and wellness benefits.
  • Equal employment opportunity and compliance-focused work environment.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →