16 часов назад
Site Reliability Engineer III (SRE) (AWS, Kubernetes)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer III (SRE) (AWS, Kubernetes): Ensuring the reliability, performance, and scalability of applications on the Guidewire Cloud Platform with an accent on automation, observability, and distributed-systems troubleshooting. Focus on designing SLI/SLO frameworks, building automated runbooks and monitoring, deploying scalable AWS and Kubernetes infrastructure, and responding to incidents through on-call rotations.
Location: Dublin, Ireland
Company
provides cloud-based software for property and casualty insurers, including core insurance, digital, analytics, and AI capabilities.
What you will do
- Partner with development teams to troubleshoot incidents and minimize customer impact.
- Develop automated runbooks, monitors, dashboards, and synthetic transactions.
- Improve application reliability, performance, scalability, and operational efficiency on the Cloud Platform.
- Deploy and manage scalable infrastructure across AWS and Kubernetes using Terraform and cloud-native approaches.
- Document incidents, identify root causes, and implement measures to prevent recurrence.
- Participate in mandatory on-call rotations, including coverage outside regular business hours, on weekends, and during holidays.
Requirements
- Professional experience in SRE or a similar reliability-focused role.
- Software engineering experience with Python, Go, or Java, including clean and testable code.
- Experience designing and implementing SLIs, SLOs, and error budgets.
- Experience with APM, telemetry, monitoring, performance optimization, and troubleshooting distributed systems on cloud infrastructure.
- Experience with CI/CD pipelines, AWS, Kubernetes, Terraform, and infrastructure configuration management through GitOps, Puppet, or Ansible.
- Understanding of cloud networking, security, and vulnerability management, including programmatic infrastructure remediation.
Nice to have
- SRE or AWS certification.
- Experience with SQL, database administration, data pipelines, performance tuning, and schema design.
- Familiarity with TeamCity, Bitbucket Pipelines, Jenkins, or GitHub Actions.
- Exposure to Hadoop, Apache Spark, Amazon Redshift, microservices, and event-driven architectures.
Culture & Benefits
- Work on cloud technology used by more than 540 insurers in 40 countries.
- Collaborate in a culture focused on innovation, teamwork, continuous learning, and work-life balance.
- Competitive compensation and comprehensive benefits.
- Opportunities for career development and skills growth.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
13 часов назад
Site Reliability Engineer Staff (Cloud Infrastructure)
15 часов назад
Senior Site Reliability Engineer (GCP)
6 дней назад
Site Reliability Engineer (Kubernetes)
84 051 - 93 390GBP
18 часов назад
Senior Site Reliability Engineer (Kubernetes)
2 дня назад
Senior Site Reliability Engineer (AWS)
2 часа назад