Senior Infrastructure SRE
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Location: Hybrid in Mississauga, Canada; candidates must reside within commuting distance of the specified office. Hybrid attendance includes regular team events at the office. Occasional travel may be required to the Mississauga and/or Salt Lake City offices.
Salary: CAD 139,000–155,000 per year, plus bonus and benefits.
Company
is a privately held health technology company providing cloud platforms, healthcare data services, and an ecosystem of more than 400 integrated partners for over 30,000 provider organizations.
What you will do
- Design and operate highly available infrastructure across Azure, AWS, GCP, and on-premises environments.
- Build reusable Infrastructure as Code with Terraform or Pulumi and establish GitOps standards.
- Automate operational workflows, including auto-remediation, self-healing systems, capacity planning, and AI-assisted investigation.
- Define SLIs, SLOs, error budgets, observability strategies, and reliability targets for critical services.
- Lead on-call response for complex infrastructure incidents, conduct blameless post-mortems, and drive systemic fixes.
- Own infrastructure domains end to end, mentor SREs, review changes, and lead multi-team reliability initiatives.
Requirements
- 5+ years of experience designing and operating cloud infrastructure, including production environments spanning multiple cloud platforms.
- Expertise with Azure or AWS, working knowledge of an additional platform, and 3+ years of production Infrastructure as Code experience with Terraform, Pulumi, or CloudFormation.
- Strong programming skills in Python, Go, or Bash, with experience writing tested and maintainable automation.
- Expertise operating Kubernetes and containerized workloads in production, plus experience with VM-based compute, service mesh, identity, SSO, storage, and cloud IAM.
- Experience applying SRE practices, defining SLIs/SLOs and error budgets, reducing MTTR and operational toil, and improving system reliability.
- Working knowledge of regulated environments such as HIPAA, SOC 2, PCI, or FedRAMP, including audit evidence, access controls, encryption, and data residency.
Nice to have
- Experience in healthcare technology or highly regulated SaaS environments.
- Cloud architecture certifications and experience operating Kubernetes across multiple clusters.
- Knowledge of Kafka, Azure Service Bus, Event Hubs, SQS, Pub/Sub, CI/CD, GitLab, GitHub Actions, or ArgoCD.
- Experience with AI-assisted engineering and operations tooling or contributions to open-source infrastructure projects.
- Bachelor’s degree in Computer Science, Computer Engineering, Information Technology, or a related field, or equivalent practical experience.
Culture & Benefits
- Transparent collaboration through OKRs, cross-functional retrospectives, and public roadmaps.
- Blameless incident management, root-cause analysis, data-driven decisions, and continuous learning.
- Flexible paid time off, wellness support, parental and caregiver leave, and fertility and adoption support.
- Retirement plan matching, continuous development support, employee assistance, and inclusion communities.
- Benefits start from Day 1.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →