21 день назад
Senior Observability Engineer (Terraform)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Observability Engineer (Terraform): Leading the development and scaling of telemetry systems for reliable, performant, and resilient platforms with an accent on metrics, traces, logs, events, and cloud-native distributed systems. Focus on building automated Terraform-based telemetry pipelines, defining SLIs and error budgets, and improving incident response through AI-assisted debugging and root cause analysis.
Location: Hybrid role requiring 2–3 days per week in the Galway, Ireland office, with 2 remote days available.
Company
operates a circular fashion platform combining subscription fashion rental, one-time rental, ownership, proprietary technology, and reverse logistics.
What you will do
- Lead the architecture, delivery, and continuous improvement of observability solutions using platforms such as Splunk Observability Cloud and Google Cloud Observability.
- Build scalable, automated telemetry pipelines with Terraform and Infrastructure-as-Code workflows.
- Define standards for metrics, traces, logs, events, instrumentation, SLIs, error budgets, and system health.
- Partner with application, platform, security, and compliance teams to integrate observability into development and operational practices.
- Improve incident response, debugging, and root cause analysis through modern AI-assisted workflows.
- Provide technical guidance, reusable frameworks, documentation, and training while participating in SRE/Platform on-call rotations.
Requirements
- 5+ years of experience in SRE, DevOps, or platform engineering, with deep expertise in observability and telemetry design.
- Expertise in observability for cloud-native, distributed systems.
- Strong knowledge of metrics, logs, traces, and events across product services and infrastructure.
- Hands-on experience with Terraform, CI/CD tooling, service instrumentation, and Kubernetes-based environments.
- Understanding of service mesh architectures, asynchronous message flows, and distributed systems behavior.
- Strong communication, ownership, collaboration, and structured root cause analysis skills.
Culture & Benefits
- Continuous integration, test-driven development, peer code reviews, pair programming, and open-source contributions.
- Paid time off, bereavement leave, and family sick leave.
- Universal paid parental leave and a flexible return-to-work program.
- Paid sabbatical after 5 years of continuous service and a stakeholder pension.
- Health, dental, and dependent care coverage from the first day of employment.
- Company-wide events and outings.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →