6 дней назад
Lead Site Reliability Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Lead Site Reliability Engineer (AI): Building an automation-first reliability ecosystem for a global SaaS platform across multi-cloud and Kubernetes environments with an accent on AI-driven operations, observability, and self-healing systems. Focus on designing predictive monitoring, automated root cause analysis, progressive delivery, and fault-tolerant infrastructure while leading incident response and mentoring engineers.
Location: Remote from Poland
Company
provides a cloud compliance platform that connects tax and technology, processing customer API calls and tax returns at global scale.
What you will do
- Own and evolve the reliability strategy for distributed SaaS systems across multi-cloud platforms.
- Design AI-driven operations, including predictive monitoring, anomaly detection, and automated root cause analysis.
- Build observability solutions with Prometheus, Grafana, and OpenTelemetry.
- Create self-healing systems and automation frameworks that reduce manual operational work.
- Improve deployment practices through feature flags, progressive delivery, and safe rollout strategies.
- Lead incident response, improve recovery times, implement lasting fixes, and mentor engineers.
Requirements
- 10+ years of experience in SaaS, distributed systems, or site reliability engineering.
- Programming experience with Go, Java, or Python.
- Deep experience with Prometheus, Grafana, and OpenTelemetry.
- Hands-on experience with Kubernetes, containerisation, and multi-cloud platforms such as AWS, GCP, Azure, or OCI.
- Strong understanding of Linux systems, networking, cloud-native architectures, automation, and reliability engineering.
- Experience applying AI or machine learning to operational workflows.
Culture & Benefits
- AI-first workflows, decision-making, and products.
- Compensation package with paid time off and paid parental leave.
- Many employees are eligible for bonuses.
- Benefits may include private medical, life, and disability insurance depending on location.
- Commitment to diversity, equity, inclusion, and employee resource groups.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
5 дней назад
Site Reliability Engineer (Cloud/Azure/SRE)
6 часов назад
Observability Engineer (Terraform)
Latitude
3 дня назад
Senior Site Reliability Engineer (Kubernetes)
5 дней назад
Principal DevOps Lead (Observability)
Т-Банк
1 день назад
Site Reliability Engineer
5 часов назад