10 часов назад
Staff Engineer, Site Reliability Engineering (Automotive)
147 000 - 196 600CAD
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Staff Engineer, Site Reliability Engineering (Automotive) (SRE/Cloud/Observability): Building reliable, observable, and scalable infrastructure for vehicle telemetry, data ingestion, and software-defined vehicle platforms with an accent on production readiness, CI/CD, and operational automation. Focus on designing fault-tolerant systems, leading incident response, developing AI-assisted workflows, and driving durable reliability improvements across engineering teams.
Location: Hybrid in Markham, Ontario; expected to report to the Markham office at least three times per week.
Salary: $147,000–$196,600 per year.
Company
develops software-defined vehicle platforms and connected mobility technologies focused on safety, emissions reduction, congestion reduction, and connected customer experiences.
What you will do
- Design and implement scalable, fault-tolerant, observable infrastructure for vehicle telemetry, data ingestion, and platform operations.
- Lead production readiness across multiple teams through hands-on coding, reliability standards, architectural improvements, and resilient deployments.
- Design and improve CI/CD pipelines with quality gates, artifact promotion, deployment verification, progressive delivery, and rollback practices.
- Define SLOs, SLIs, monitoring, alerting, runbooks, and operational practices with SRE, product, application, infrastructure, and data engineering teams.
- Build reusable AI workflows and evaluations, automate incident operations, and develop diagnostics, remediation, and customer-facing status workflows.
- Lead on-call incident response, post-incident reviews, system-level fixes, technical direction, mentoring, and cross-functional reliability initiatives.
Requirements
- 8+ years of experience in SRE, DevOps, or systems engineering, including managing or mentoring high-impact teams.
- Experience designing and operating high-scale, cloud-native production systems, preferably on Azure, AWS, or GCP.
- Hands-on expertise in observability, standardized instrumentation, OpenTelemetry collectors, SLO/SLI definitions, monitors, alerts, and dashboards.
- Experience with production readiness, service ownership, incident management, on-call rotations, post-incident learning, and continuous reliability improvement.
- Experience with CI/CD, GitOps, release strategies, quality gates, progressive delivery, deployment verification, and safe rollback.
- Strong programming ability in Python, Go, Java, or a comparable language, plus familiarity with AI-assisted development, LLM applications, agentic workflows, and evaluation techniques.
Nice to have
- Azure Databricks, Azure Event Hubs, Azure Kubernetes Service, Helm, Kustomize, or Terraform experience.
- Experience with GitHub Actions, Argo CD, Prometheus, Grafana, Datadog, or OpenTelemetry platforms.
- Experience with Promptfoo, Copilot-based AI skills, or LLM application development and testing.
- Experience operating large-scale systems using Fivetran, Apache Flink, Kafka, or Pulsar.
- Experience with vehicle telemetry, connected-vehicle platforms, or high-volume event-driven systems.
Culture & Benefits
- Paid time off, vacation days, holidays, and supplemental pregnancy, parental, and adoption leave benefits.
- Healthcare, dental, vision, and life insurance benefits.
- Defined Contribution Pension plan with company and matching contributions.
- GM Vehicle Purchase Plan for employees and their families.
- GM does not provide immigration-related sponsorship for this role.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →