2 дня назад
Senior Site Reliability Engineer
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Site Reliability Engineer (AWS/Kubernetes): Building and operating reliable multi-tenant SaaS infrastructure for Guidewire’s cloud platform and InsuranceSuite products with an accent on automation, observability, and production-scale reliability. Focus on developing self-healing systems, improving incident management, and supporting containerized microservices across 24x7 operations.
Location: Ireland — Dublin. The role includes rotating weekend on-call support for 24x7 customer operations, primarily aligned with PST timings, and occasional travel to other offices.
Company
provides cloud software, analytics, digital products, and AI-enabled platforms for property and casualty insurance companies worldwide.
What you will do
- Engineer and operate resilient multi-tenant SaaS infrastructure and customer-focused application environments.
- Automate AWS infrastructure, deployments, operational tooling, and recurring tasks.
- Develop reliability improvements and tooling for containerized microservices and core infrastructure.
- Build observability tooling, metrics, dashboards, and self-healing capabilities.
- Improve incident management through risk identification, mitigation, SLO tracking, and blameless postmortems.
- Collaborate with product engineering teams, contribute code, and create documentation and training materials.
Requirements
- Bachelor’s degree in Computer Science or a related field.
- Strong software engineering and automation skills with Bash, Python, and/or Go.
- Deep Linux engineering experience and extensive AWS automation experience.
- Hands-on experience with Docker, Helm, Kubernetes/EKS, CNI, and ingress networking in production.
- Experience with Terraform or related IaC tools, GitOps and DevOps tooling, microservices, and production-scale operations.
- Experience with SSO technologies, SAML, OAuth, x.509 certificates, relational databases, and observability tools such as Datadog, CloudWatch, and PagerDuty.
Nice to have
- Experience with Okta, Kafka, AWS SQS, KubeVela, or Crossplane.
- Experience supporting Java, Apache, and Tomcat web applications in production.
- Experience applying AI and data-driven insights to improve productivity and reliability.
Culture & Benefits
- Collaborative culture focused on integrity, rationality, collegiality, curiosity, and innovation.
- Mentoring, teaching, and cross-functional collaboration are part of the role.
- Rotating weekend operational support is required for customer emergencies.
- Occasional travel of less than 5% may be required for training and team meetings.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
2 дня назад
Senior Site Reliability Engineer (Kubernetes)
7 дней назад
Senior Infrastructure Engineer
1 день назад
Site Reliability Engineer (AWS)
5 часов назад
Senior Site Reliability Engineer (Kubernetes)
12 часов назад
Senior Software DevOps Engineer (Medtech)
65 600 - 98 400€
2 дня назад