5 дней назад
Senior Site Reliability Engineer
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Site Reliability Engineer (Cloud Infrastructure/Observability): Building and improving cloud infrastructure, software delivery, and observability tooling for distributed, fault-tolerant systems with an accent on Kubernetes, Terraform, CI/CD, and platform reliability. Focus on automating operations, developing alerting and self-healing capabilities, coordinating production incident response, and leading infrastructure initiatives at scale.
Location: Hybrid in Galway, Ireland; 2–3 days per week in the office and 2 days remotely
Company
operates a circular fashion platform combining designer apparel rentals, proprietary technology, and reverse logistics.
What you will do
- Lead technology initiatives across cloud infrastructure, software delivery, and observability.
- Build tooling, policies, and processes that improve platform scale, performance, and reliability.
- Develop observability, alerting, tracing, automation, and self-healing capabilities.
- Coordinate platform operations, incident response, issue reporting, escalation, and remediation.
- Automate maintenance and operations through CI/CD, self-service, and developer tooling.
- Lead assigned projects and promote Site Reliability Engineering practices across development and operations teams.
Requirements
- At least 5 years of hands-on experience with orchestration tools such as Kubernetes.
- Advanced experience with Terraform, Helm, or Ansible, plus CI/CD tools such as GitHub and Artifactory.
- Experience with monitoring, alerting, and logging tools, including Splunk or GCP Monitoring.
- At least 3 years maintaining production environments on GCP, AWS, or Azure.
- At least 5 years developing and delivering products with Bash, Python, Golang, or Java.
- Willingness to participate in on-call rotations, troubleshoot production issues, and conduct root-cause analyses.
Culture & Benefits
- Pragmatic, entrepreneurial engineering environment with continuous integration, test-driven development, peer reviews, and pair programming.
- Paid annual leave, bereavement leave, and family sick leave.
- Universal paid parental leave and a flexible return-to-work program.
- Paid sabbatical after five years of continuous service.
- Stakeholder pension and health, dental, and dependent care coverage from the first day of employment.
- Company-wide events and outings.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →