1 час назад
Site Reliability Engineer (AI Platform)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer (AI Platform): Keeping production platform systems reliable for ML and LLM workloads across public cloud environments with an accent on infrastructure automation, observability, secure-by-default operations, and model-serving reliability. Focus on building GPU-backed inference infrastructure, monitoring model performance and drift, creating reusable GitOps and Terraform tooling, and managing SLOs, incident response, cost, and compliance tradeoffs.
Location: Hybrid in Reading, England, United Kingdom, or Berlin, Germany
Company
combines advanced technology with a global network of people to make unusable data usable and support AI-driven real-world impact.
What you will do
- Own the reliability of the AI Platform, including GPU-backed model-serving and inference infrastructure for ML and LLM workloads.
- Define and operate SLOs, on-call processes, and incident response for production models and services.
- Build observability for services and ML models, including drift monitoring, LLM tracing, evaluations, guardrails, metrics, and logging.
- Shape company-wide platform direction and create golden paths, reusable developer tooling, GitHub Actions, GitOps workflows, and Terraform modules.
- Package reusable infrastructure components such as Grafana, Istio, CloudNative tools, model registries, and feature stores.
- Embed security, compliance, cost governance, data-residency controls, PII protection, and audit trails into the platform.
Requirements
- 5+ years of experience in infrastructure engineering, DevOps, or SRE operating large-scale, highly available production systems with Kubernetes.
- Hands-on production experience with Kubernetes, Helm, Terraform or CloudFormation, and at least one major cloud provider; AWS is preferred.
- Good proficiency in Python, Go, or scripting for automation and tooling.
- Experience owning at least one infrastructure build end to end with measurable outcomes such as deployment time, MTTR, cost, adoption, or availability.
- Strong first-principles problem solving, cross-functional collaboration, and communication across global teams and time zones.
- Willingness to support 24x7 operational processes.
Nice to have
- Experience running ML workloads on Kubernetes, including GPU scheduling, capacity, and cost management.
- Production-scale model serving with KServe, Ray Serve, Triton, vLLM, or similar tools.
- Experience with MLOps tooling such as Kubeflow, MLflow, Feast, or Weights & Biases.
- Production LLMOps experience covering inference serving, prompt and version management, tracing, evaluations, drift, guardrails, and cost per request.
- Experience governing ML and LLM workloads with data-residency, PII, and audit requirements.
Culture & Benefits
- Mission-driven environment focused on economic and social impact.
- People-centric workplace supporting growth, well-being, belonging, and authentic self-expression.
- Global collaboration across diverse cultures and perspectives.
- Opportunities for professional development, continuous learning, and meaningful contribution.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
DeepL
6 дней назад
Senior Platform Engineer (Kubernetes)
5 дней назад
Staff Platform Site Reliability Engineer (Kubernetes)
6 дней назад
Site Reliability Engineer (Cloud Banking)
5 дней назад
DevOps Engineer (.NET/C#)
6 дней назад
Site Reliability Engineer
3 500 - 4 500€
3 дня назад