Назад
Company hidden
5 дней назад

Lead Site Reliability Engineer (AWS/GCP)

154 000 - 200 000CAD
Формат работы
remote (только Canada)
Тип работы
fulltime
Грейд
lead
Английский
b2
Страна
US/Canada
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Lead Site Reliability Engineer (AWS/GCP): Designing and evolving a multi-cloud, multi-region active-active content-serving platform handling more than 25 billion requests per day with an accent on infrastructure automation, observability, distributed systems, and reliability strategy. Focus on scaling the platform toward 50 billion daily requests, architecting Kubernetes and logging platforms, establishing SLOs, and leading complex incident response and performance improvements.

Location: Ontario, Canada (Remote)

Salary: CAD 154,000–200,000 per year, plus potential bonus and benefits

Company

hirify.global scales content personalization for marketers through data-activated content generation and AI decisioning.

What you will do

  • Define and drive infrastructure automation strategies that reduce manual work and improve performance and incident outcomes.
  • Own the architecture, reliability, and evolution of core platform applications in a multi-cloud, multi-region active-active environment.
  • Architect the logging platform while balancing availability, retention, and cost optimization.
  • Establish capacity planning and performance management frameworks and guide complex troubleshooting.
  • Lead cross-functional reliability initiatives with SRE and service engineering teams.
  • Mentor engineers and identify systemic platform weaknesses and improvement opportunities autonomously.

Requirements

  • 6+ years of hands-on experience in Site Reliability or Software Engineering, including leading multi-cloud architecture and strategy across AWS and GCP.
  • Experience designing and operating scalable, resilient distributed systems, including Apache Pulsar, Apache Kafka, Grafana Loki, and ScyllaDB/Cassandra.
  • Experience leading observability platforms and defining observability standards and SLO frameworks using Prometheus, Thanos, Grafana Alloy, Loki, and Tempo.
  • Expertise in infrastructure as code with Terraform and Chef, plus advanced Kubernetes expertise across EKS and GKE.
  • Proficiency in NodeJS, Golang, Ruby, Python, and shell scripting, with advanced Linux systems expertise.
  • Experience improving on-call, monitoring, alerting, automated runbooks, incident response, diagnostics, and performance tuning.

Culture & Benefits

  • Remote work arrangement for Ontario, Canada.
  • Full medical, financial, and other benefits are available.
  • Every SRE team member participates in a week-long on-call rotation.
  • Inclusive and equal-opportunity workplace committed to supporting diverse employees.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →