2 дня назад
Senior Site Reliability Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Site Reliability Engineer (AI/GCP): Building automated validation, progressive delivery, observability, and recovery systems for an AI-assisted platform that optimizes enterprise Google Cloud costs with an accent on production reliability, deployment safety, and infrastructure reproducibility. Focus on defining SLOs and error budgets, developing agent-driven operational workflows, improving GCP infrastructure, and turning incidents into durable safeguards and automation.
Location: Jakarta, Indonesia
Company
is a Google Cloud Partner delivering infrastructure, data analytics, machine learning, cloud migration, and application development solutions, including the Rabbit cloud cost optimization product.
What you will do
- Build automated validation, progressive rollout, and recovery mechanisms for AI-assisted delivery.
- Define and operatione SLOs, SLIs, and error budgets to balance delivery speed and reliability.
- Develop AI agent workflows for alert triage, incident investigation, maintenance, and reporting, with verification and human approval where needed.
- Improve logging, metrics, tracing, alerting, runbooks, incident investigations, and blameless postmortems.
- Extend Terraform and delivery tooling to keep infrastructure reproducible and changes reviewable.
- Strengthen GCP infrastructure, networking, access controls, capacity management, performance, reliability, and cost efficiency.
Requirements
- 6+ years of experience in SRE, production engineering, or infrastructure-heavy backend roles with ownership of production systems.
- Hands-on Google Cloud experience, including Cloud Run, networking, and IAM; GCP expertise will be assessed during the interview.
- Strong Terraform, CI/CD, deployment safety, observability, troubleshooting, and root-cause analysis skills.
- Coding ability in Go, Python, or a comparable language, with experience building maintainable operational tooling.
- Experience leading production incident investigations and implementing improvements that prevent recurrence.
- Strong written English and the ability to collaborate asynchronously with a distributed team.
Nice to have
- Kubernetes or GKE experience, including deployment, operation, debugging, and scaling of containerized services.
- Datadog experience with dashboards, monitors, logs, APM, and distributed tracing.
- Experience with progressive delivery, policy-as-code, automated rollback, GCP cost management, FinOps, or security.
Culture & Benefits
- AI-assisted engineering and agent-driven workflows are part of the standard development process.
- Work on systems with direct customer impact in enterprise Google Cloud environments.
- Small-team environment with short decision paths and end-to-end ownership.
- Engineering work includes infrastructure, delivery, and operations automation with production safeguards.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
7 дней назад
Site Reliability Engineer III
8 дней назад
Site Reliability Engineer (Cloud Banking)
8 дней назад
Sr. Site Reliability Engineer
160 000 - 180 000$
3 дня назад
Senior/Lead Site Reliability Engineer (AI/LLM)
6 дней назад
Site Reliability Engineer – Quantum Computing (f/m/d)
7 дней назад