Staff Site Reliability Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
TL;DR
Staff Site Reliability Engineer (AI): Designing and optimizing the reliability, performance, and cost of a scalable AI platform with an accent on AI Gateway architecture and LLM provider integrations. Focus on building observability for non-deterministic systems, managing inference latency, and implementing FinOps strategies for AI workloads.
Location: Hybrid in Amsterdam, Netherlands. Relocation support for the employee and their family is provided.
Company
provides AI-powered automation tools for creators and brands to engage audiences across social platforms like Instagram, Messenger, and WhatsApp.
What you will do
- Own the reliability and performance of AI infrastructure, including AI Gateway and integrations with Amazon Bedrock and Azure OpenAI.
- Design and evolve the AI Gateway's routing, failover, rate limiting, caching, and guardrails.
- Build observability for AI systems, defining SLOs for latency, throughput, and quality per model and provider.
- Drive FinOps and cost optimization for AI workloads through model right-sizing and caching strategies.
- Manage capacity planning and incident response for inference services, including runbooks and postmortems.
- Scale AI expertise across the organization by setting technical standards and coaching teams on LLM features.
Requirements
- 5+ years of experience in SRE, platform, or infrastructure engineering at significant scale.
- Hands-on experience operating LLM-backed systems in production (e.g., OpenAI, Anthropic, Bedrock).
- Deep cloud-native background with AWS, Kubernetes, Terraform/IaC, and CI/CD.
- Strong observability practice using Prometheus, Grafana, or OpenTelemetry.
- Proven track record of implementing cloud and inference cost-optimization.
- Must be based in or be able to relocate to the Netherlands.
Nice to have
- Experience building or operating LLM gateways/proxies such as LiteLLM or Kong AI Gateway.
- Experience with GPU workload optimization, quantization, or serving frameworks (vLLM, TGI, Triton).
- Experience with evaluation pipelines and quality monitoring for LLM outputs.
Culture & Benefits
- Relocation support for you and your family.
- Comprehensive health insurance and a flexible benefits budget for wellbeing and home office.
- Professional development budget for conference tickets and online courses.
- Hybrid work model with generous, flexible time off.
- In-office perks including free meals, snacks, and company-funded sport activities.
- AI-first work culture with default access to tools like Claude and OpenAI.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →