обновлено 1 день назад
Site Reliability Engineer (Kubernetes)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer (Kubernetes) (Cloud Infrastructure): Building and operating highly available, fault-tolerant cloud infrastructure for a platform processing large volumes of real-time customer data with an accent on Kubernetes, Infrastructure as Code, and observability. Focus on automating operational work, troubleshooting production systems, improving deployment reliability, and designing systems that scale without single points of failure.
Location: Remote, Riyadh, Saudi Arabia
Company
is an AI-native customer experience intelligence platform that manages customer lifecycles using proprietary multilingual NLU technology.
What you will do
- Design and maintain highly available, fault-tolerant, and scalable infrastructure.
- Manage and optimize cloud workloads across AWS, GCP, or Azure using Terraform and other Infrastructure as Code tools.
- Operate, troubleshoot, scale, and upgrade production Kubernetes clusters and containerized workloads.
- Implement monitoring and alerting with tools such as Prometheus, Grafana, Datadog, or ELK.
- Respond to incidents, lead root cause analysis, and improve reliability based on failure learnings.
- Automate operational work and collaborate with DevOps and engineering teams on CI/CD, performance, and deployment reliability.
Requirements
- Approximately 3 years of experience in SRE, DevOps, or infrastructure engineering.
- Hands-on production experience with Kubernetes, Docker, and cloud environments such as AWS, GCP, or Azure.
- Experience with Terraform or similar Infrastructure as Code tools.
- Ability to write automation scripts in Python, Bash, or similar languages.
- Understanding of CI/CD pipelines, networking, load balancing, distributed systems, and high-availability design.
- Experience implementing actionable monitoring and alerting with Prometheus, Grafana, Datadog, or ELK.
Nice to have
- Production experience with RabbitMQ or Redis.
- Familiarity with Ansible or AWX.
- Experience with multi-cloud or hybrid environments.
- AWS, GCP, or Linux certifications, or an ITI background.
Culture & Benefits
- Remote work arrangement.
- Ownership of infrastructure reliability and continuous improvement.
- Focus on reducing manual work through automation.
- Collaboration with DevOps and engineering teams to improve the wider system.
Hiring process
- Talent Acquisition screening interview.
- Technical interview with the SRE Lead, followed by a technical task.
- Final interview with the SRE Lead and Cloud DevOps Director.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
7 дней назад
Senior DevOps Engineer (Cloud)
155 000 - 175 000$
2 дня назад
DevOps Engineer (AI)
163 000 - 204 000$
15 часов назад
Senior DevOps / Platform Reliability Engineer (AI Platform)
4 дня назад
AWS Cloud Infrastructure Engineer (AI)
4 дня назад
Staff Site Reliability Engineer (Linux/Network Troubleshooting/Scripting) (AI)
4 часа назад