2 дня назад
Senior Site Reliability Engineer (Kubernetes)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Site Reliability Engineer (Kubernetes): Improving the reliability, resilience, security, and availability of production systems with an accent on Kubernetes, cloud infrastructure, automation, and observability. Focus on operating live infrastructure, responding to incidents, reducing toil, and developing resilient solutions across distributed teams.
Location: Remote, limited to Texas, the United States of America, Alberta, or British Columbia
Company
provides Network as a Service that connects businesses to cloud providers, data centers, and each other.
What you will do
- Improve production reliability, system resilience, security, and availability within an SRE team.
- Operate live production infrastructure, handle alerts, participate in on-call rotations, and respond to incidents.
- Develop automation, write code and effective runbooks, reduce operational toil, and prevent recurring problems.
- Use observability systems for metrics, logs, and traces while maintaining a strong signal-to-noise ratio.
- Collaborate with engineering teams and stakeholders on requirements, demonstrations, peer reviews, and technical solutions.
- Conduct blameless post-incident reviews and promote DevOps, SRE, and industry best practices.
Requirements
- 5+ years administering Linux systems and related infrastructure in production.
- Strong Kubernetes and cloud infrastructure fundamentals; AWS experience is strongly preferred.
- Experience with Bash and either Python or Go, infrastructure as code with Terraform, CI/CD, version control, and GitHub.
- Experience with at least one of PostgreSQL, Cassandra, or ClickHouse and with production observability stacks.
- Knowledge of SRE practices including SLIs, SLOs, SLAs, error budgets, blast radius, and blameless postmortems.
- Self-directed, collaborative work style suited to an asynchronous, globally distributed team.
Nice to have
- Bare-metal infrastructure experience.
- AWS experience and Terraform expertise.
Culture & Benefits
- Remote-first working environment with coworking options.
- Four weeks of paid annual leave, parental leave, birthday leave, and purchased annual leave options.
- Wellness allowance and employee wellbeing initiatives.
- Study and training allowance plus five days of paid study leave.
- Inclusive, collaborative environment with recognition programs and modern workspaces.
Hiring process
- Candidates who meet the selection criteria are invited to an interview.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
2 дня назад
Site Reliability Engineer (Kubernetes)
180 000 - 220 000$
3 дня назад
Senior Site Reliability Engineer (AWS/Kubernetes)
60 000 - 85 500€
8 дней назад
Site Reliability Engineer (Kubernetes)
2 дня назад
Senior Site Reliability Engineer (Kubernetes)
170 000 - 185 000$
5 дней назад
Senior Site Reliability Engineer
6 дней назад
Staff Site Reliability Engineer (AI)
252 000 - 308 000$