6 часов назад
Senior Site Reliability Engineer (Kubernetes)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Site Reliability Engineer (Kubernetes): Maintaining and scaling reliable Kubernetes clusters and Hydrolix deployments across multiple cloud platforms with an accent on infrastructure reliability, CI/CD, observability, and incident response. Focus on automating operations, analyzing root causes of failures, optimizing distributed systems, and supporting customers through on-call coverage.
Location: Remote, APAC
Company
provides a cloud data platform built for petabyte-scale datasets, helping organizations reduce data costs while increasing data retention.
What you will do
- Deploy, maintain, and improve reliable Kubernetes clusters and deployments across multiple cloud platforms.
- Design and optimize systems for service reliability, availability, performance, and operational efficiency.
- Build and maintain CI/CD tools and processes for efficient, dependable deployments.
- Develop monitoring, alerting, and incident response strategies, and conduct root cause analyses after failures.
- Automate repetitive operational tasks and implement long-term preventive measures.
- Collaborate with engineering, infrastructure, product, and customer teams while participating in weekday business-hours and once-monthly weekend on-call coverage.
Requirements
- Bachelor’s degree in Computer Science, Engineering, or a related field.
- At least five years of experience supporting complex distributed systems as an SRE or in a similar role.
- Experience with observability and debugging tools such as Prometheus, Vector, Grafana, Superset, or Kibana.
- Proficiency with at least one major cloud platform: AWS, GCP, Azure, or Linode.
- Experience with SQL databases and strong Linux administration, performance tuning, and system-level troubleshooting skills.
- Proficiency in Python, Go, or Rust, plus strong written and verbal communication skills for working with customers and cross-functional teams.
Culture & Benefits
- Remote collaboration within a distributed engineering team.
- Global team coordination to provide round-the-clock support.
- Hands-on work focused on operational excellence and SRE best practices.
- Direct customer engagement to investigate and resolve incidents.
Hiring process
- Submit the application form with a resume.
- Cover letter and relevant references may be required depending on the role.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
4 дня назад
Senior Site Reliability Engineer (Kubernetes)
145 000 - 193 000CAD
6 дней назад
Senior Site Reliability Engineer (Kubernetes)
3 дня назад
Senior Cloud Site Reliability Engineer
8 часов назад
Staff Site Reliability Engineer (Linux/Network Troubleshooting/Scripting) (AI)
1 день назад
Engineer III, Site Reliability (SRE)
4 дня назад
Site Reliability Engineer (Kubernetes)
84 051 - 93 390GBP