11 дней назад
Senior Site Reliability Engineer - Platform Reliability (Resilience) (Kubernetes)
62 800 - 84 300€
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Site Reliability Engineer - Platform Reliability (Resilience) (Kubernetes): Building and scaling a multi-cloud platform for Elastic Cloud Hosted and Serverless services with an accent on infrastructure automation, Kubernetes operations, and system resilience. Focus on preventing customer impact, improving incident management and observability, and developing reliable tooling for distributed systems at scale.
Location: Portugal
Base salary: €62,800–€84,300 per year. The role has no variable compensation component.
Company
develops the Search AI Platform and cloud-based solutions for search, security, and observability.
What you will do
- Design, build, scale, and mature the multi-cloud platform supporting Cloud Hosted and Serverless services.
- Develop software, tooling, and automation for platform infrastructure and system engineering.
- Lead initiatives that improve the reliability and scalability of ’s global infrastructure.
- Respond to major incidents, prevent repeated customer impact, and improve problem-management practices.
- Participate in a follow-the-sun on-call rotation, primarily during working hours.
- Collaborate with distributed engineering teams through inclusive communication, coaching, and mentoring.
Requirements
- Software engineering experience with a customer-first approach to operational problems and site reliability engineering.
- Experience operating SaaS products in public cloud environments, preferably using Infrastructure-as-Code tools such as Terraform or Crossplane.
- Experience building or operating Kubernetes infrastructure at scale across multiple cloud providers.
- Ability to write non-trivial programs in Golang or other programming languages and work with containerized services such as Docker.
- Professional Linux system administration experience on large-scale distributed systems.
- Experience with alerting, metrics, major incident management, and the Stack or comparable observability systems such as Prometheus, Graphite, or Influx.
Nice to have
- Experience with public cloud and managed Kubernetes services.
- Experience working remotely or in globally distributed teams.
Culture & Benefits
- Inclusive, collaborative environment focused on operational excellence and continuous improvement.
- Flexible locations and schedules for many roles.
- Health coverage for employees and families in many locations.
- Generous vacation allowance and at least 16 weeks of parental leave.
- Up to $2,000 or local-currency equivalent in matched charitable donations and up to 40 volunteer hours annually.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →