12 часов назад
Staff Site Reliability Engineer (AWS/Kubernetes)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Staff Site Reliability Engineer (AWS/Kubernetes): Building and evolving reliable, scalable AWS infrastructure, CI systems, and developer tooling for a platform supporting more than 9 million users with an accent on Terraform-managed infrastructure, production Kubernetes, observability, and incident response. Focus on improving EKS, databases, deployment reliability, actionable alerting, and platform architecture across the engineering organization.
Location: Remote-first; hiring in Canada
Company
is a product technology company serving expecting and new families through ecommerce, health, media, mobile, and related products.
What you will do
- Own and evolve AWS infrastructure with Terraform, including EKS clusters, databases, and core services.
- Improve the speed and reliability of CI/CD systems used across the engineering organization.
- Support developers across local, staging, and production environments.
- Establish actionable monitoring and alerting standards using observability tools.
- Lead or support incident response, post-incident reviews, and corrective actions.
- Contribute to platform strategy and architectural decisions for long-term infrastructure evolution.
Requirements
- Deep hands-on Terraform and infrastructure-as-code expertise.
- Proven AWS experience at scale, including EKS, RDS, cloud networking, DNS, CDNs, and load balancers.
- Production Kubernetes operations experience and the ability to troubleshoot complex issues.
- Experience designing and improving CI/CD systems such as CircleCI or GitHub Actions.
- Strong observability, on-call, and incident management experience.
- Regular use of AI tools to improve engineering speed and output.
Culture & Benefits
- Remote-first work with team members across the U.S. and Canada.
- Twice-yearly company gatherings.
- Company-paid medical, dental, and vision insurance.
- Retirement savings plan with company matching, flexible spending accounts, PTO, and paid parental leave.
- Paid company-wide Winter Wonder Week and a remote work stipend.
- Health, parenting, childcare, and financial planning benefits.
Hiring process
- Interviews are recorded and transcribed using an interview recording tool.
- AI is used to support application review and assessment, while hiring decisions are made by people.
- Candidates are expected to represent their own thinking accurately during the application and interview process.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
2 дня назад
Staff Site Reliability Engineer (AWS/Kubernetes)
140 000 - 155 000CAD
18 часов назад
Principal Site Reliability Engineer (AWS/Kubernetes)
163 620 - 212 710$
18 часов назад
Senior Site Reliability Engineer (Kubernetes)
12 часов назад
Site Reliability Engineer (AWS/Kubernetes)
200 000 - 300 000$
Latitude
6 дней назад
Senior Site Reliability Engineer (Kubernetes)
9 минут назад