2 дня назад
Site Reliability Engineer
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer (Kubernetes/AWS): Maintaining and improving the global Tyk Cloud platform with an accent on reliability engineering, infrastructure automation, observability, and incident management. Focus on operating multi-region and multi-cloud Kubernetes environments, building metrics and dashboards, automating operations, and resolving complex production reliability issues.
Location: Remote; listed hiring locations include Argentina, Brazil, Canada, Colombia, Costa Rica, and Mexico. Remote working from anywhere in the world is offered. On-call coverage is required from 16:00–04:00 UTC.
Company
provides an API Management platform that helps organizations connect systems and services through its cloud and B2B products.
What you will do
- Maintain and improve the global Cloud platform within defined service-level objectives.
- Identify and resolve reliability issues with the engineering squad.
- Expand the platform’s multi-region and multi-cloud capabilities.
- Build metrics, dashboards, monitoring, and logging systems.
- Participate in on-call rotation, incident management, post-incident analysis, and penetration-testing support.
- Automate operational tasks, document SRE processes, and improve operational efficiency and running costs.
Requirements
- Experience launching and operating production-scale Kubernetes clusters and containerized infrastructure.
- Advanced AWS/EKS and Linux administration experience, plus infrastructure design on AWS or other cloud providers.
- Proficiency with Terraform, infrastructure as code, and Helm.
- Experience operating MongoDB or other document databases, Redis or other key-value stores, and distributed software.
- Experience with Prometheus, Grafana, Thanos, logging systems, networking concepts, and common networking protocols.
- Strong collaboration skills and willingness to participate in on-call rotation from 16:00–04:00 UTC.
Nice to have
- Experience with GCP or Azure, bare-metal infrastructure, API management, or large-scale distributed storage.
- Familiarity with Rancher and CKA, CKAD, or CKS certifications.
- Experience creating and delivering production software in Go.
Culture & Benefits
- Unlimited paid holiday and flexible working hours.
- Remote-first work with a distributed international team.
- Employee share scheme.
- Generous maternity and paternity leave.
- Company retreats and a culture centered on autonomy, responsibility, inclusion, experimentation, and continuous improvement.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →