4 дня назад
Site Reliability Engineer (Kubernetes)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer (Kubernetes): Building and operating cloud-native infrastructure for distributed database systems on Kubernetes, AWS, and Google Cloud with an accent on automation, observability, scalability, and reliability. Focus on writing Golang automation, optimizing CI/CD pipelines, troubleshooting complex infrastructure issues, and designing disaster recovery and high-availability solutions.
Location: Remote within Europe, preferably in the European time zone
Company
develops a unified multimodel contextual data platform that connects enterprise data with LLMs, copilots, and AI agents.
What you will do
- Design, implement, and maintain cloud infrastructure on AWS and Google Cloud.
- Operate and improve the scalability, performance, and reliability of Kubernetes-based distributed database systems.
- Write production-grade Golang code with development teams to automate infrastructure management and operations.
- Optimize CI/CD pipelines, deployment processes, monitoring, logging, and alerting systems.
- Develop disaster recovery, high-availability, and fault-tolerance strategies.
- Troubleshoot network, operating system, cloud infrastructure, and customer-impacting issues, including participation in on-call rotations.
Requirements
- Proven experience as an SRE or DevOps Engineer in a cloud-native environment.
- Strong Kubernetes experience managing large-scale distributed systems.
- Experience with AWS and Google Cloud, networking, security practices, and Linux internals.
- Familiarity with Docker, CI/CD tools, Git, and monitoring and observability tools such as Prometheus, Grafana, and ELK.
- Programming experience with Golang or Python, plus strong troubleshooting, communication, and collaboration skills.
- Ability to self-organize and work independently as part of a remote team.
Nice to have
- Experience with distributed databases or large-scale data storage systems.
- Experience with Python or Bash scripting, Terraform, GitOps, and cloud security best practices.
- Strong Golang programming skills for developing automation tools, scripts, or services.
Culture & Benefits
- Work on AI and data infrastructure for enterprise applications.
- Collaborate with experienced engineering, marketing, and product teams.
- Help shape the contextual data layer for AI-powered systems.
- Support critical production systems through on-call rotations.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
6 дней назад
Site Reliability Engineer (Kubernetes)
6 дней назад
Staff Software Engineer, Compute Platform (Kubernetes)
7 дней назад
Senior Site Reliability Engineer (AWS)
109 000 - 128 000€
6 дней назад
Senior Site Reliability Engineer (Cloud Infrastructure)
150 000 - 172 000$
7 дней назад
Senior Site Reliability Engineer (Azure)
1 день назад
Senior Manager of Site Reliability Engineering (AWS)
155 000 - 170 000$