7 дней назад
Senior Site Reliability Engineer (Kubernetes)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Site Reliability Engineer (Kubernetes): Designing, building, scaling, and maturing Elastic’s multi-cloud platform for hosted and serverless services with an accent on infrastructure automation, Kubernetes operations, and platform resilience. Focus on preventing customer impact, improving incident management, developing reliability tooling, and operating distributed systems through a follow-the-sun on-call rotation.
Location: United Kingdom
Company
develops the Search AI Platform for search, security, and observability across structured and unstructured data.
What you will do
- Design, build, scale, and mature the multi-cloud platform hosting Cloud Hosted and Serverless services.
- Develop software, tooling, and automation that support platform infrastructure and product deployments.
- Lead technical initiatives to automate system engineering and improve global infrastructure reliability.
- Respond to major incidents, prevent repeated customer impact, and improve problem-management practices.
- Participate in a follow-the-sun on-call rotation, primarily during working hours.
- Promote collaboration, operational excellence, inclusive communication, coaching, and mentoring.
Requirements
- Experience operating SaaS products in public cloud environments using infrastructure-as-code tools such as Crossplane or Terraform.
- Experience building or operating Kubernetes infrastructure at scale across multiple cloud providers and automating its operation.
- Ability to write non-trivial programs in Golang or another programming language and work with containerized services such as Docker.
- Professional Linux system administration experience on large-scale distributed systems.
- Experience improving alerting, metrics, and major incident management using tools such as Stack, Graphite, Prometheus, or Influx.
- Experience designing and implementing solutions with the Stack and working effectively in globally distributed teams.
Nice to have
- Experience with public cloud and managed Kubernetes services.
- Experience working remotely or in distributed teams.
Culture & Benefits
- Distributed work environment focused on inclusion, collaboration, and progress over perfection.
- Competitive pay based on the work performed rather than previous salary.
- Health coverage for employees and families in many locations.
- Flexible locations and schedules for many roles, plus generous vacation time.
- Parental leave of at least 16 weeks, charitable donation matching up to $2,000, and up to 40 volunteer hours annually.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →