Назад
Company hidden
8 дней назад

Senior Site Reliability Engineer (SRE & AI Platform Operations)

Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Spain/Netherlands
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Site Reliability Engineer (SRE & AI Platform Operations) (GCP/Terraform/AI): Owning production reliability, observability, deployment safety, cloud cost optimization, and AI runtime infrastructure across high-velocity microservices and Google Cloud workloads with an accent on SLOs, progressive delivery, FinOps, and operationalized LLM integrations. Focus on designing automated rollbacks and health gates, controlling model and token budgets, validating disaster recovery, and eliminating operational toil through infrastructure automation and guardrails.

Location: Rotterdam, Netherlands; work from the Rotterdam office five days a week

Company

hirify.global is rebuilding its platform with a new frontend, pricing engine, content engine, and product engine for an AI-driven, high-growth business.

What you will do

  • Define and enforce SLOs, SLIs, and error-budget policies across checkout, payments, catalog pipelines, routing engines, and other critical services.
  • Expand distributed observability and diagnostic tooling using Telemetry, Sentry, Google Cloud Monitoring, and tracing across microservices, Laravel Horizon workers, Redis, and Google Cloud.
  • Improve CI/CD with canary traffic shifting, automated SLO-driven rollbacks, and health gates in GitHub Actions.
  • Operate AI agents, semantic pipelines, and background automation while managing latency, rate limits, queue backpressure, provider availability, and model and token budgets.
  • Optimize Google Cloud infrastructure costs and lead incident response, post-mortems, capacity planning, resilience, and disaster recovery validation.
  • Own Terraform and Google Cloud Run infrastructure, IAM least privilege, Secret Manager, runbooks, and self-service deployment tooling.

Requirements

  • Production SRE or platform experience operating high-traffic distributed systems with strict uptime, latency, and release-safety requirements.
  • Strong troubleshooting skills across Linux, containers, relational and NoSQL databases, Redis queues, and cloud network boundaries.
  • Extensive hands-on experience with Google Cloud Platform, Google Cloud Run, and Terraform.
  • Experience with cloud and AI runtime cost monitoring, analysis, and optimization.
  • Practical experience with Sentry, Google Cloud Monitoring, distributed tracing, SLI tracking, and canary gates.
  • Proficiency in Python or TypeScript/JavaScript, plus working familiarity with PHP in a Laravel ecosystem and operationalizing LLM integrations or AI pipelines.

Culture & Benefits

  • International environment with more than 100 professionals from over 20 nationalities.
  • Ownership from day one, experienced leadership, and freedom to drive projects.
  • Opportunities to grow within the role and across hirify.global.
  • 24/7 HelloFit gym access and an UrbanSportsClub discount.
  • Company-sponsored healthy meals and company events.

Hiring process

  • Submit a CV or resume and optionally a cover letter.
  • Answer questions about reducing cloud or AI infrastructure waste, operating LLMs in production, and implementing automation or architectural guardrails.
  • Confirm availability for five days per week in the Rotterdam office and clarify work authorization or sponsorship needs for the Netherlands.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →