Senior Site Reliability Engineer (remote within EMEA)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
TL;DR
Senior Site Reliability Engineer (AWS/Kubernetes): Building and operating highly available cloud infrastructure and self-service platform capabilities for global on-demand commerce with an accent on Kubernetes, infrastructure as code, observability, and security. Focus on architecting large-scale automation, solving complex production reliability issues, optimizing cloud costs, and setting engineering standards across teams.
Location: Remote within EMEA, with the option to work from the Riga office.
Company
is part of FYUL, a global on-demand commerce platform formed through the merger of , Printify, and Snow Commerce.
What you will do
- Architect and operate secure, scalable infrastructure across AWS accounts and environments using Terraform and related infrastructure-as-code practices.
- Design and manage production EKS clusters, cloud networking, databases, messaging systems, and platform services.
- Drive automation and GitOps adoption with Terraform, Terragrunt, and ArgoCD to improve consistency and reduce manual operations.
- Develop observability and reliability practices using Grafana, Prometheus, Loki, Tempo, and Mimir.
- Lead on-call incident response, write runbooks, ADRs, and postmortems, and improve production resilience.
- Mentor SREs, establish technical standards, collaborate with product engineering teams, and drive security and FinOps initiatives.
Requirements
- Several years of hands-on production infrastructure or SRE experience at a senior individual-contributor level.
- Strong Linux administration and Python scripting skills, plus extensive AWS experience with EKS, IAM, VPC, RDS, S3, and SQS.
- Production Kubernetes expertise, including Helm, CNI networking, Cilium, IPAM, scaling, and container security.
- Proficiency with Terraform, including modules and state management, plus GitOps experience with ArgoCD.
- Experience with production databases such as PostgreSQL, MySQL, MongoDB, or Aurora, and with Jenkins or GitHub Actions CI/CD.
- Practical incident management, observability-stack maintenance, strong written communication, initiative ownership, and mentoring experience.
Nice to have
- GCP experience.
- Kafka or AWS MSK experience.
- Experience in regulated or compliance-sensitive environments.
- Experience contributing to a platform or developer-experience roadmap used by other engineering teams.
Culture & Benefits
- Global, inclusive, and collaborative working environment.
- Flexible remote work or office work in Riga, with flexible start times up to 11 AM.
- Private health insurance and additional paid wellness and celebration days.
- Internal and external learning opportunities, mentorship, meetups, and hackathons.
- Office lunch in Riga, employee merchandise discounts, and team-building events.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →