6 дней назад
Senior Site Reliability Engineer (SRE) - Guidewire Cloud Platform (Application)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Site Reliability Engineer (SRE) - Guidewire Cloud Platform (Application) (AWS/Kubernetes): Building and operating reliable, scalable applications on the Guidewire Cloud Platform with an accent on automation, observability, and performance optimization. Focus on designing SLI/SLO and error-budget practices, engineering CI/CD pipelines and Datadog monitoring, and debugging distributed systems during on-call operations.
Location: Dublin, Ireland
Company
provides cloud software, analytics, digital products, and AI capabilities used by property and casualty insurers worldwide.
What you will do
- Build and evolve the SRE practice for applications running on the Cloud Platform.
- Collaborate with development teams to troubleshoot distributed systems and reduce customer impact.
- Develop automated runbooks, improve operational automation, and eliminate toil.
- Monitor, maintain, and optimize application reliability, performance, scalability, and availability.
- Design and improve SLI, SLO, error-budget, monitoring, dashboard, and synthetic-transaction capabilities.
- Participate in mandatory on-call rotations, including incidents outside business hours, weekends, and holidays.
Requirements
- Senior-level SRE experience with a record of improving system reliability.
- Strong software engineering skills in Python, Go, or Java, including clean and testable code.
- Experience designing and deploying SLIs, SLOs, error budgets, APM and telemetry solutions.
- Experience triaging and debugging distributed systems on cloud infrastructure.
- Experience engineering CI/CD pipelines across Kubernetes and legacy environments.
- Experience with Datadog, AWS, Kubernetes, Terraform, GitOps, Puppet, or Ansible, plus cloud networking, security, and vulnerability management.
Nice to have
- Multiple SRE or AWS certifications.
- Experience with distributed-system architectures, microservices, and event-driven systems.
- Proficiency with SQL, database administration, data pipelines, performance tuning, and schema design.
- Experience with TeamCity, Bitbucket Pipelines, Jenkins, GitHub Actions, Hadoop, Apache Spark, or Amazon Redshift.
Culture & Benefits
- Mission-driven work supporting insurance customers during crises and other difficult events.
- Culture focused on innovation, teamwork, continuous learning, and work-life balance.
- Competitive compensation and comprehensive benefits.
- Career development opportunities and collaboration with experienced technical peers.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →