Назад
Company hidden
6 часов назад

Senior Site Reliability Engineer (Kubernetes)

Формат работы
remote (только APAC)
Тип работы
fulltime
Грейд
senior
Английский
b2
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Site Reliability Engineer (Kubernetes): Maintaining and scaling reliable Kubernetes clusters and Hydrolix deployments across multiple cloud platforms with an accent on infrastructure reliability, CI/CD, observability, and incident response. Focus on automating operations, analyzing root causes of failures, optimizing distributed systems, and supporting customers through on-call coverage.

Location: Remote, APAC

Company

hirify.global provides a cloud data platform built for petabyte-scale datasets, helping organizations reduce data costs while increasing data retention.

What you will do

  • Deploy, maintain, and improve reliable Kubernetes clusters and hirify.global deployments across multiple cloud platforms.
  • Design and optimize systems for service reliability, availability, performance, and operational efficiency.
  • Build and maintain CI/CD tools and processes for efficient, dependable deployments.
  • Develop monitoring, alerting, and incident response strategies, and conduct root cause analyses after failures.
  • Automate repetitive operational tasks and implement long-term preventive measures.
  • Collaborate with engineering, infrastructure, product, and customer teams while participating in weekday business-hours and once-monthly weekend on-call coverage.

Requirements

  • Bachelor’s degree in Computer Science, Engineering, or a related field.
  • At least five years of experience supporting complex distributed systems as an SRE or in a similar role.
  • Experience with observability and debugging tools such as Prometheus, Vector, Grafana, Superset, or Kibana.
  • Proficiency with at least one major cloud platform: AWS, GCP, Azure, or Linode.
  • Experience with SQL databases and strong Linux administration, performance tuning, and system-level troubleshooting skills.
  • Proficiency in Python, Go, or Rust, plus strong written and verbal communication skills for working with customers and cross-functional teams.

Culture & Benefits

  • Remote collaboration within a distributed engineering team.
  • Global team coordination to provide round-the-clock support.
  • Hands-on work focused on operational excellence and SRE best practices.
  • Direct customer engagement to investigate and resolve incidents.

Hiring process

  • Submit the application form with a resume.
  • Cover letter and relevant references may be required depending on the role.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →