обновлено 1 месяц назад
Site Reliability Engineer (AI)
55 000 - 68 000€
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer (AI/Cloud): Building and operating reliable, scalable, and fault-tolerant cloud systems for an AI-powered personal and entrepreneurial resource planner with an accent on observability, infrastructure automation, and high availability. Focus on designing disaster recovery and failover strategies, improving CI/CD pipelines, optimizing cloud performance and costs, and leading incident response through on-call rotations.
Location: Lisbon, Portugal; on-site
Salary: €55K–€68K annually, to be discussed
Company
is a self-funded software company founded in Lisbon in 2018 that develops an AI-powered Personal & Entrepreneurial Resource Planner and has offices in Lisbon and San Francisco.
What you will do
- Design and implement scalable, reliable, and fault-tolerant systems across cloud environments.
- Develop observability solutions for monitoring, logging, and alerting using tools such as Prometheus, Grafana, Datadog, and ELK.
- Automate infrastructure provisioning, deployments, and incident response with Infrastructure as Code.
- Improve system performance, scalability, CI/CD pipelines, and incident response workflows.
- Design and maintain load balancing, failover, high-availability, and disaster recovery strategies.
- Collaborate with development and DevOps teams, conduct root cause analysis, and implement preventative measures.
Requirements
- 4+ years of experience in Site Reliability Engineering, DevOps, or Systems Engineering.
- Strong knowledge of AWS, Azure, or Google Cloud Platform and cloud-native architectures.
- Experience with observability tools, including Prometheus, Grafana, ELK, Datadog, or New Relic.
- Proficiency with Terraform, CloudFormation, or Pulumi and hands-on experience with Docker, Kubernetes, and Helm.
- Strong Linux administration and networking fundamentals, plus scripting skills in Bash, Python, or Go.
- Experience with incident management, debugging, distributed systems, security best practices, access control, and compliance.
Culture & Benefits
- Apple hardware ecosystem for work.
- Annual bonus and top-tier health and life insurance.
- Transportation budget, Coverflex benefits, childcare support, and pension fund.
- Air Conference for team collaboration and professional growth.
- Urban Sports Club membership and free meals at the hub.
- Inclusive workplace welcoming diverse backgrounds, experiences, and perspectives.
Hiring process
- Applicants must submit their own work without AI-generated assistance.
- Use of AI in application materials, assessments, or interviews results in disqualification.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
14 дней назад
Senior SRE (Site Reliability Engineer) – Modernized Application Operations
145 000 - 170 000$
13 дней назад
Senior Site Reliability Engineer (Healthcare)
200 000 - 240 000$
ClickHouse
1 день назад
Infrastructure Engineer (AWS)
90 000 - 160 000€
13 дней назад
Site Reliability Engineer (AI)
200 000 - 400 000$
12 дней назад
Senior Cloud Infrastructure and Networking
125 000 - 135 000$
13 дней назад
Senior Manager, Site Reliability Engineering (AI Ops)
222 000 - 300 500$