8 дней назад
Senior Platform Engineer — AI Infrastructure
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Platform Engineer — AI Infrastructure (GCP/Terraform): Owning production reliability, observability, deployment safety, cloud cost optimization, and AI runtime infrastructure across high-velocity microservices and cloud workloads with an accent on SLOs, progressive delivery, FinOps, and operational guardrails. Focus on designing resilient serverless platforms, automating SLO-driven rollbacks, controlling model and token costs, and validating disaster recovery against strict RTO/RPO targets.
Location: Rotterdam, Netherlands; onsite five days a week
Company
is rebuilding its platform with new frontend, pricing, content, and product engines for an AI-driven, high-growth business.
What you will do
- Own production reliability through SLOs, SLIs, error budgets, incident response, post-mortems, and automated safeguards.
- Expand distributed observability and diagnostic tooling across microservices, Laravel Horizon workers, Redis, and Google Cloud infrastructure.
- Improve CI/CD with canary traffic shifting, automated health gates, and SLO-driven rollbacks in GitHub Actions.
- Architect and operate AI agent runtimes, semantic pipelines, and background automation while managing latency, rate limits, queue backpressure, provider availability, and token costs.
- Drive FinOps, capacity planning, load testing, dependency isolation, and disaster recovery validation against strict RTO/RPO targets.
- Own Terraform and Google Cloud Run infrastructure, IAM security, secrets management, internal tooling, runbooks, and self-service deployment primitives.
Requirements
- Production SRE or platform engineering experience with high-traffic distributed systems and safe release cycles.
- Strong troubleshooting skills across Linux, containers, relational and NoSQL databases, Redis queues, and cloud networking.
- Hands-on experience with Google Cloud Platform, Google Cloud Run, and Terraform.
- Experience optimizing cloud and AI runtime costs while balancing performance, reliability, and budgets.
- Practical experience with Sentry, Google Cloud Monitoring, distributed tracing, SLI tracking, and canary gates.
- Proficiency in Python or TypeScript/JavaScript, with working familiarity with modern PHP and Laravel; experience operationalizing LLM integrations, embeddings, vector workflows, or background task orchestration.
Culture & Benefits
- International environment with more than 100 professionals representing over 20 nationalities.
- Ownership from day one, experienced leadership, and freedom to drive technical projects.
- Opportunities to grow within the role and across .
- 24/7 gym access, UrbanSportsClub discount, company events, and company-sponsored healthy meals.
Hiring process
- Application includes questions about cloud and AI cost optimization, production LLM operations, and automation or architectural guardrails.
- Applicants must confirm that working from the Rotterdam office five days a week is suitable.
- Applicants must indicate their Netherlands visa sponsorship status.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
12 дней назад
Staff Software Engineer (AI/DevX)
11 дней назад
Senior Platform Engineer (Containers)
11 дней назад
Senior Platform Engineer (Agentic AI & Harness Engineering)
11 дней назад
Senior Platform Engineer
150 000 - 200 000$
10 дней назад
Lead Research Platform Engineer (AI)
5 дней назад