Назад
Company hidden
2 дня назад

Senior Site Reliability Engineer (AI)

Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Singapore/Indonesia/Italy +2 еще
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Site Reliability Engineer (AI/GCP): Building automated validation, progressive delivery, observability, and recovery systems for an AI-assisted platform that optimizes enterprise Google Cloud costs with an accent on production reliability, deployment safety, and infrastructure reproducibility. Focus on defining SLOs and error budgets, developing agent-driven operational workflows, improving GCP infrastructure, and turning incidents into durable safeguards and automation.

Location: Jakarta, Indonesia

Company

hirify.global is a Google Cloud Partner delivering infrastructure, data analytics, machine learning, cloud migration, and application development solutions, including the Rabbit cloud cost optimization product.

What you will do

  • Build automated validation, progressive rollout, and recovery mechanisms for AI-assisted delivery.
  • Define and operationhirify.globale SLOs, SLIs, and error budgets to balance delivery speed and reliability.
  • Develop AI agent workflows for alert triage, incident investigation, maintenance, and reporting, with verification and human approval where needed.
  • Improve logging, metrics, tracing, alerting, runbooks, incident investigations, and blameless postmortems.
  • Extend Terraform and delivery tooling to keep infrastructure reproducible and changes reviewable.
  • Strengthen GCP infrastructure, networking, access controls, capacity management, performance, reliability, and cost efficiency.

Requirements

  • 6+ years of experience in SRE, production engineering, or infrastructure-heavy backend roles with ownership of production systems.
  • Hands-on Google Cloud experience, including Cloud Run, networking, and IAM; GCP expertise will be assessed during the interview.
  • Strong Terraform, CI/CD, deployment safety, observability, troubleshooting, and root-cause analysis skills.
  • Coding ability in Go, Python, or a comparable language, with experience building maintainable operational tooling.
  • Experience leading production incident investigations and implementing improvements that prevent recurrence.
  • Strong written English and the ability to collaborate asynchronously with a distributed team.

Nice to have

  • Kubernetes or GKE experience, including deployment, operation, debugging, and scaling of containerized services.
  • Datadog experience with dashboards, monitors, logs, APM, and distributed tracing.
  • Experience with progressive delivery, policy-as-code, automated rollback, GCP cost management, FinOps, or security.

Culture & Benefits

  • AI-assisted engineering and agent-driven workflows are part of the standard development process.
  • Work on systems with direct customer impact in enterprise Google Cloud environments.
  • Small-team environment with short decision paths and end-to-end ownership.
  • Engineering work includes infrastructure, delivery, and operations automation with production safeguards.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →