Назад
Company hidden
обновлено 1 день назад

Site Reliability Engineer (Kubernetes)

Формат работы
remote (только Saudi_arabia)
Тип работы
fulltime
Грейд
middle
Английский
b2
Страна
SA
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Site Reliability Engineer (Kubernetes) (Cloud Infrastructure): Building and operating highly available, fault-tolerant cloud infrastructure for a platform processing large volumes of real-time customer data with an accent on Kubernetes, Infrastructure as Code, and observability. Focus on automating operational work, troubleshooting production systems, improving deployment reliability, and designing systems that scale without single points of failure.

Location: Remote, Riyadh, Saudi Arabia

Company

hirify.global is an AI-native customer experience intelligence platform that manages customer lifecycles using proprietary multilingual NLU technology.

What you will do

  • Design and maintain highly available, fault-tolerant, and scalable infrastructure.
  • Manage and optimize cloud workloads across AWS, GCP, or Azure using Terraform and other Infrastructure as Code tools.
  • Operate, troubleshoot, scale, and upgrade production Kubernetes clusters and containerized workloads.
  • Implement monitoring and alerting with tools such as Prometheus, Grafana, Datadog, or ELK.
  • Respond to incidents, lead root cause analysis, and improve reliability based on failure learnings.
  • Automate operational work and collaborate with DevOps and engineering teams on CI/CD, performance, and deployment reliability.

Requirements

  • Approximately 3 years of experience in SRE, DevOps, or infrastructure engineering.
  • Hands-on production experience with Kubernetes, Docker, and cloud environments such as AWS, GCP, or Azure.
  • Experience with Terraform or similar Infrastructure as Code tools.
  • Ability to write automation scripts in Python, Bash, or similar languages.
  • Understanding of CI/CD pipelines, networking, load balancing, distributed systems, and high-availability design.
  • Experience implementing actionable monitoring and alerting with Prometheus, Grafana, Datadog, or ELK.

Nice to have

  • Production experience with RabbitMQ or Redis.
  • Familiarity with Ansible or AWX.
  • Experience with multi-cloud or hybrid environments.
  • AWS, GCP, or Linux certifications, or an ITI background.

Culture & Benefits

  • Remote work arrangement.
  • Ownership of infrastructure reliability and continuous improvement.
  • Focus on reducing manual work through automation.
  • Collaboration with DevOps and engineering teams to improve the wider system.

Hiring process

  • Talent Acquisition screening interview.
  • Technical interview with the SRE Lead, followed by a technical task.
  • Final interview with the SRE Lead and Cloud DevOps Director.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →