Назад
Company hidden
1 час назад

Site Reliability Engineer (AI Platform)

Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
UK/Germany
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Site Reliability Engineer (AI Platform): Keeping production platform systems reliable for ML and LLM workloads across public cloud environments with an accent on infrastructure automation, observability, secure-by-default operations, and model-serving reliability. Focus on building GPU-backed inference infrastructure, monitoring model performance and drift, creating reusable GitOps and Terraform tooling, and managing SLOs, incident response, cost, and compliance tradeoffs.

Location: Hybrid in Reading, England, United Kingdom, or Berlin, Germany

Company

hirify.global combines advanced technology with a global network of people to make unusable data usable and support AI-driven real-world impact.

What you will do

  • Own the reliability of the AI Platform, including GPU-backed model-serving and inference infrastructure for ML and LLM workloads.
  • Define and operate SLOs, on-call processes, and incident response for production models and services.
  • Build observability for services and ML models, including drift monitoring, LLM tracing, evaluations, guardrails, metrics, and logging.
  • Shape company-wide platform direction and create golden paths, reusable developer tooling, GitHub Actions, GitOps workflows, and Terraform modules.
  • Package reusable infrastructure components such as Grafana, Istio, CloudNative tools, model registries, and feature stores.
  • Embed security, compliance, cost governance, data-residency controls, PII protection, and audit trails into the platform.

Requirements

  • 5+ years of experience in infrastructure engineering, DevOps, or SRE operating large-scale, highly available production systems with Kubernetes.
  • Hands-on production experience with Kubernetes, Helm, Terraform or CloudFormation, and at least one major cloud provider; AWS is preferred.
  • Good proficiency in Python, Go, or scripting for automation and tooling.
  • Experience owning at least one infrastructure build end to end with measurable outcomes such as deployment time, MTTR, cost, adoption, or availability.
  • Strong first-principles problem solving, cross-functional collaboration, and communication across global teams and time zones.
  • Willingness to support 24x7 operational processes.

Nice to have

  • Experience running ML workloads on Kubernetes, including GPU scheduling, capacity, and cost management.
  • Production-scale model serving with KServe, Ray Serve, Triton, vLLM, or similar tools.
  • Experience with MLOps tooling such as Kubeflow, MLflow, Feast, or Weights & Biases.
  • Production LLMOps experience covering inference serving, prompt and version management, tracing, evaluations, drift, guardrails, and cost per request.
  • Experience governing ML and LLM workloads with data-residency, PII, and audit requirements.

Culture & Benefits

  • Mission-driven environment focused on economic and social impact.
  • People-centric workplace supporting growth, well-being, belonging, and authentic self-expression.
  • Global collaboration across diverse cultures and perspectives.
  • Opportunities for professional development, continuous learning, and meaningful contribution.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →