Site Reliability Engineer (AWS)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
TL;DR
Site Reliability Engineer (AWS): Operating and evolving highly available AWS cloud infrastructure with an accent on SLOs, incident response, infrastructure as code, and observability. Focus on automating toil, designing fault-tolerant systems, improving deployment safety, and optimizing production reliability at scale.
Location: Remote within Portugal; candidates must be based in Portugal and have the legal right to work in Portugal or the wider European Union.
Company
provides technology and engineering services for cloud and infrastructure environments.
What you will do
- Define and enforce SLOs, SLIs, and error budgets to guide engineering priorities and release decisions.
- Own incident detection, response, mitigation, post-mortems, and on-call improvements.
- Design and maintain highly available AWS architectures across compute, networking, storage, and data layers.
- Build infrastructure as code and CI/CD pipelines for repeatable and low-risk deployments.
- Automate operational toil and develop self-healing systems, observability, alerting, and runbooks.
- Partner with development and security teams on capacity, performance, cost optimization, production readiness, security, and compliance.
Requirements
- Based in Portugal and legally authorized to work in Portugal or the wider European Union.
- Fluent English.
- Production experience operating AWS systems at scale, including EC2, ECS/EKS, Lambda, VPC, IAM, S3, RDS, and CloudWatch.
- Strong SRE knowledge covering SLOs, error budgets, toil reduction, and blameless post-mortems.
- Hands-on experience with Terraform or other infrastructure-as-code tools, Kubernetes/EKS, and CI/CD.
- Strong scripting and automation skills in Python, Go, or Bash, plus experience with observability, networking, Linux, distributed systems, incident response, and on-call operations.
Nice to have
- AWS certifications such as Solutions Architect, DevOps Engineer, or SysOps.
- Experience with multi-account or multi-region AWS environments and cost governance.
- Chaos engineering, resilience testing, service mesh, GitOps, or policy-as-code experience.
- Background in regulated industries such as financial services or healthcare.
Culture & Benefits
- Fully remote work arrangement within Portugal.
- Hands-on ownership of cloud reliability and operational health.
- Blameless incident management and continuous improvement of on-call practices.
- Collaboration with development and security teams on production engineering.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →