5 дней назад
Senior Site Reliability Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Site Reliability Engineer (AI): Building reliable, scalable cloud and Kubernetes infrastructure, healthy PostgreSQL data platforms, and AI platform capabilities for an e-commerce operations platform with an accent on observability, infrastructure as code, data performance, and secure AI delivery. Focus on owning SLOs and incident response, designing GitOps-based platforms, tuning database replication and capacity, and building guardrails for production AI agents.
Location: Berlin, Germany
Total compensation: €90,000–€120,000 depending on experience; on-call compensation is paid separately.
Company
is building an operations platform that helps brands manage e-commerce operations through technology and a network of operations partners.
What you will do
- Own production reliability through SLOs, alerting, observability, incident response, on-call participation, and post-incident improvements.
- Design, build, and optimize AWS and Kubernetes infrastructure managed with Terraform and delivered through GitOps.
- Operate the PostgreSQL fleet, including performance, capacity, replication, upgrades, backup, and recovery.
- Build AI templates, data access models, coding standards, and guardrails for production AI agents.
- Embed reliability into product delivery and make targeted changes in Rails and React codebases when appropriate.
- Harden cloud and Kubernetes security through least-privilege access, secrets management, access controls, and vulnerability remediation.
Requirements
- Senior-level experience running production systems on AWS and Kubernetes, ideally Amazon EKS.
- Hands-on infrastructure-as-code experience with Terraform.
- Production PostgreSQL experience covering query and index tuning, replication, upgrades, and backup and recovery.
- Experience with observability, SLOs, incident response, and on-call operations.
- Ability to write production code in Python, Ruby, TypeScript, or a similar language.
- Experience using AI coding tools and shipping an LLM-backed product or capability beyond a chat interface.
Nice to have
- Ruby on Rails experience.
- GitOps with Argo CD or a similar tool.
- Experience with data pipelines, cloud data warehouses, or advanced cloud security.
Culture & Benefits
- Collaborative culture focused on trust, empowerment, constructive feedback, and professional growth.
- Virtual employee stock options for full-time employees and hardware selected according to personal preference.
- Thirty vacation days annually, with a sabbatical opportunity after three years.
- Monthly wellness and productivity budget, flexible working hours, and additional technical equipment.
- Shared on-call rotation with additional compensation and regular team events.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →