1 день назад
Site Reliability Expert (Observability)
100 000 - 150 000CAD
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Expert (Observability): Defining and implementing observability strategies, monitoring solutions, and SRE practices across cloud-native microservices environments with an accent on SLIs, SLOs, error budgets, distributed tracing, and operational governance. Focus on troubleshooting complex production issues, automating infrastructure and operations, and improving reliability with engineering and product teams.
Location: Remote in Canada
Salary: $100,000–$150,000 CAD annually, depending on experience and location.
Company
is an experience innovation company helping recognized brands create digital products and transform customer experiences through data, creativity, and technology.
What you will do
- Define and implement observability strategies, standards, and governance across applications and platforms.
- Design and maintain monitoring, alerting, dashboards, and reporting using Dynatrace or equivalent platforms.
- Establish SRE practices including SLIs, SLOs, error budgets, and symptom-based alerting.
- Partner with engineering and product teams to improve reliability, performance, and operational maturity.
- Analyze distributed systems and troubleshoot complex production issues using monitoring and tracing data.
- Lead technical workstreams, coach teams, and promote documentation and continuous improvement.
Requirements
- Significant SRE experience in large-scale production environments.
- Deep knowledge of SLIs, SLOs, error budgets, symptom-based alerting, APM, RUM, instrumentation, RBAC, tagging, and governance.
- Experience with OpenTelemetry, distributed tracing, microservices architectures, AWS, and Kubernetes.
- Practical experience with Terraform, Bash, Python, and CI/CD tools such as GitLab CI.
- Ability to lead technical initiatives, manage stakeholders, and work autonomously in complex operational and agile environments.
- Excellent communication skills in both French and English are required.
Nice to have
- Experience monitoring Java Spring Boot applications.
- Experience with e-commerce platforms and high-transaction environments.
- Experience establishing enterprise-wide observability frameworks and governance models.
- Consulting or advisory experience supporting multiple engineering teams.
Culture & Benefits
- Continuous learning, professional growth, creativity, autonomy, and knowledge-sharing.
- Comprehensive insurance options with employer contributions of up to 80%, including short- and long-term disability coverage.
- Virtual healthcare, employee and family assistance, and mental health support.
- $500 personal spending account and $30 monthly personal technology reimbursement.
- RRSP matching through a Deferred Profit Sharing Plan, up to 4%.
- Flexible vacation, winter holiday closure, and flexible scheduling.
Hiring process
- The Talent Acquisition team reviews applications and contacts candidates whose experience aligns with the role.
- Applications should include relevant experience and expertise; candidate evaluation is based on skills, experience, and potential.
- Reasonable accommodations are available during the interview process.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
3 дня назад
Senior DevOps Developer (AWS)
107 000 - 157 300CAD
4 дня назад
Senior DevOps Developer
107 000 - 157 300$
4 дня назад
Sr. Site Reliability Engineer (Healthcare)
125 000 - 145 000$
4 дня назад
Senior Site Reliability Engineer (GovCloud)
117 000 - 209 330$
7 дней назад
Sr. Site Reliability Engineer (AWS/CI/CD)
2 дня назад
Platform Engineer (AWS)
170 000 - 210 000$