2 дня назад
Site Reliability Team Leader
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Team Leader (Cloud Infrastructure/SRE): Leading a global SRE team while building reliable, scalable, and highly available production systems with an accent on Kubernetes, cloud infrastructure, automation, observability, and deployment processes. Focus on designing reliability initiatives, guiding incident response, improving operational excellence, and developing SRE engineers across Israel and Ukraine.
Location: Israel-based, Tel Aviv; leading a global SRE team with engineers in Israel and Ukraine.
Company
develops an AI-powered Positionless Marketing platform that helps marketers analyze, create, launch, and optimize personalized campaigns.
What you will do
- Lead, mentor, hire, onboard, and support the career development and performance management of a global SRE team.
- Set the technical roadmap for reliability engineering, automation, observability, and operational excellence.
- Design and guide improvements to the reliability, availability, and scalability of the production cloud platform.
- Oversee internal tooling, automation, CI/CD, Kubernetes infrastructure, and safe production rollouts using Canary, Blue/Green, and Feature Flag strategies.
- Drive monitoring, alerting, dashboards, incident response, root cause analysis, and long-term preventive improvements.
- Partner with Software Engineering, DevOps, DBA, and Product leadership on engineering priorities.
Requirements
- 5+ years of experience in Site Reliability Engineering, Platform Engineering, DevOps, or Infrastructure Engineering, including experience leading or mentoring engineers.
- Hands-on production experience with Kubernetes and public cloud platforms such as GCP or AWS.
- Strong programming and scripting skills; Python is preferred.
- Experience building automation and internal engineering tools, working with CI/CD pipelines, and applying modern deployment methodologies.
- Experience with observability platforms such as Datadog, Prometheus, or Grafana, plus strong knowledge of Linux, networking, distributed systems, and cloud-native architectures.
- Fluent English, written and spoken, is required for collaboration across multiple countries and time zones.
Nice to have
- Formal people management or team lead experience.
- Infrastructure as Code experience with Terraform or Ansible.
- Experience with Kafka, Pub/Sub, Redis, OpenTelemetry, SRE principles, SLIs, SLOs, and error budgets.
- Experience with large-scale SaaS production environments or relevant cloud and Kubernetes certifications.
Culture & Benefits
- SRE operates as an engineering discipline focused on automation, reliability, and software delivery.
- Work with large-scale cloud infrastructure, Kubernetes, distributed systems, and enterprise SaaS workloads.
- Collaborate with a global SRE organization across multiple locations.
- Lead reliability initiatives that improve customer experience and engineering productivity.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
13 часов назад
Engineering Manager (Site Reliability)
6 дней назад
Head of Engineering
240 000 - 270 000$
8 часов назад
Platform Engineering Manager (AWS)
140 000 - 175 000$
3 часа назад
Engineering Manager - Developer Experience (m/f/d)
Morpho Labs
6 дней назад
Technical Lead Manager Infrastructure (Web3)
GB
5 часов назад