5 дней назад
Lead Site Reliability Engineer (AWS/GCP)
154 000 - 200 000CAD
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Lead Site Reliability Engineer (AWS/GCP): Designing and evolving a multi-cloud, multi-region active-active content-serving platform handling more than 25 billion requests per day with an accent on infrastructure automation, observability, distributed systems, and reliability strategy. Focus on scaling the platform toward 50 billion daily requests, architecting Kubernetes and logging platforms, establishing SLOs, and leading complex incident response and performance improvements.
Location: Ontario, Canada (Remote)
Salary: CAD 154,000–200,000 per year, plus potential bonus and benefits
Company
scales content personalization for marketers through data-activated content generation and AI decisioning.
What you will do
- Define and drive infrastructure automation strategies that reduce manual work and improve performance and incident outcomes.
- Own the architecture, reliability, and evolution of core platform applications in a multi-cloud, multi-region active-active environment.
- Architect the logging platform while balancing availability, retention, and cost optimization.
- Establish capacity planning and performance management frameworks and guide complex troubleshooting.
- Lead cross-functional reliability initiatives with SRE and service engineering teams.
- Mentor engineers and identify systemic platform weaknesses and improvement opportunities autonomously.
Requirements
- 6+ years of hands-on experience in Site Reliability or Software Engineering, including leading multi-cloud architecture and strategy across AWS and GCP.
- Experience designing and operating scalable, resilient distributed systems, including Apache Pulsar, Apache Kafka, Grafana Loki, and ScyllaDB/Cassandra.
- Experience leading observability platforms and defining observability standards and SLO frameworks using Prometheus, Thanos, Grafana Alloy, Loki, and Tempo.
- Expertise in infrastructure as code with Terraform and Chef, plus advanced Kubernetes expertise across EKS and GKE.
- Proficiency in NodeJS, Golang, Ruby, Python, and shell scripting, with advanced Linux systems expertise.
- Experience improving on-call, monitoring, alerting, automated runbooks, incident response, diagnostics, and performance tuning.
Culture & Benefits
- Remote work arrangement for Ontario, Canada.
- Full medical, financial, and other benefits are available.
- Every SRE team member participates in a week-long on-call rotation.
- Inclusive and equal-opportunity workplace committed to supporting diverse employees.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
Mercury
2 дня назад
Software Engineer - Infrastructure (AWS)
122 400 - 158 400$
5 дней назад
Senior Site Reliability Engineer (Cybersecurity)
131 000 - 164 250CAD
5 дней назад
Lead Infrastructure Engineer
110 000 - 186 000$
6 дней назад
Staff Site Reliability Engineer (AWS/Kubernetes)
140 000 - 155 000CAD
5 дней назад
Lead DevOps Engineer
150 000 - 170 000$
7 дней назад
Senior IT Cloud Operations Engineering Specialist (GCP)
104 400 - 135 700CAD