обновлено 5 дней назад
Staff Software Reliability Engineer (Data Platform)
160 000 - 220 000CAD
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Staff Software Reliability Engineer (Data Platform) (Streaming Data/ML Infrastructure): Designing and operating high-volume, low-latency distributed data-platform services that power analytics and machine learning with an accent on scalable streaming infrastructure, data quality, and observability. Focus on building high-performance platform components, tuning distributed systems, debugging production issues, and managing reliability across complex service stacks.
Location: Toronto, Ontario, Canada; hybrid
Salary: CAD 160,000–220,000 annually, plus equity where applicable, bonus, and benefits.
Company
Okta builds identity and security infrastructure that helps organizations securely use cloud, data, and AI technologies.
What you will do
- Design, implement, deploy, and own high-performance, scalable data-platform components.
- Build and optimize streaming infrastructure for analytics and machine-learning systems.
- Collaborate with engineers, architects, and cross-functional partners on system design and implementation.
- Conduct design reviews, code reviews, technical analysis, and performance tuning.
- Coach and mentor engineers while helping scale the engineering organization.
- Debug production issues, participate in on-call rotation, and manage incidents across services and stack layers.
Requirements
- 5+ years of industry experience.
- 2+ years of experience with an object-oriented language, preferably Java.
- Hands-on experience with cloud-based distributed computing technologies, including messaging, data processing, storage, compute, coordination, and scheduling systems.
- Experience developing and tuning highly scalable distributed systems.
- Strong understanding of software engineering principles, multithreading, garbage collection, and memory management.
- Experience with reliability engineering, including data quality, data observability, and incident management.
Nice to have
- Experience with security, encryption, identity management, or authentication infrastructure.
- Experience building mission-critical, high-volume services with major public cloud providers.
- Experience developing data integration applications for petabyte-scale batch and online systems.
- Experience with distributed systems such as Kafka or Hadoop.
- Experience developing Kubernetes-based services on AWS.
Culture & Benefits
- Opportunity to work on foundational data services supporting streaming analytics, reporting, machine learning, and product telemetry.
- Ownership of technically challenging projects using modern data-platform technologies.
- Health, dental, and vision insurance.
- RRSP matching, healthcare spending support, telemedicine, and paid leave including PTO and parental leave.
- Equity where applicable, bonus opportunities, and an in-person onboarding experience.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
12 дней назад
Jr Software Engineer (SRE)
70 000 - 80 000CAD
10 дней назад
Senior Operations Reliability Engineer - IAM (AI)
92 000 - 120 800$
5 дней назад
Senior SRE (AI SaaS)
100 000 - 180 000$
9 дней назад
Site Reliability Engineer (Fintech)
6 дней назад
Senior Site Reliability Engineer - Batch Management (FinTech)
12 дней назад
Staff Engineer, Site Reliability Engineering (Automotive)
147 000 - 196 600CAD