Назад
Company hidden
6 дней назад

Platform Engineer - LLM Inference Infrastructure (Go/Kubernetes)

Формат работы
remote (только Europe)/hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
UK/US/Europe +1 еще
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Platform Engineer - LLM Inference Infrastructure (Go/Kubernetes): Building backend services and a Kubernetes-native, multi-tenant platform for LLM inference with an accent on usage metering, billing correctness, access control, and production reliability. Focus on writing Go services, custom controllers, Helm charts, network policies, and observability components while proving correctness in production.

Location: Based in Helsinki or London, or remote in Europe; hybrid work mode

Company

hirify.global builds a full-stack AI cloud spanning data centers, hardware, and a cloud platform for AI teams.

What you will do

  • Build and maintain production Go services for the platform API, customer usage API, metering, and billing.
  • Develop the Kubernetes-native platform through GitOps, custom resources, controllers, progressive rollouts, network policies, and certificate management.
  • Implement multi-tenancy, tenant provisioning, project-level model access, and credential handling across environments.
  • Build usage metering and billing reconciliation that accurately accounts for per-tenant token usage.
  • Improve observability with logs, metrics, and traces that support incident investigation.

Requirements

  • 4+ years of professional experience.
  • Strong production experience with Go, including concurrency, context propagation, and error handling.
  • Deep Kubernetes knowledge, including Helm charts, network policies, RBAC, and control-loop behavior.
  • Strong SQL and relational data-modelling skills with PostgreSQL, including transactions, indexes, and migration safety.
  • Ability to verify changes in production and communicate testing limits clearly.
  • Clear written English for design notes, runbooks, and incident write-ups.

Nice to have

  • Experience building Kubernetes operators and controllers with CRDs, controller-runtime, Kubebuilder, or Operator SDK.
  • GitOps at scale with Argo CD or Flux, including multi-cluster environments.
  • LLM serving experience with vLLM, SGLang, TensorRT-LLM, KServe, GPU scheduling, batching, or KV cache behavior.
  • Experience with distributed messaging, ingress, time-series or log stores, frontend maintenance, billing systems, or Python.

Culture & Benefits

  • Cash and equity compensation.
  • Healthcare, lunch, wellbeing, and other fringe benefits.
  • Profitable operations with rapid, sustained growth.
  • Collaboration with engineers, researchers, and partners across the global AI ecosystem.

Hiring process

  • Applications are submitted through the Careers page; applications by email are not accepted.
  • The position has no artificial application deadline, and hiring proceeds once a suitable candidate is found.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →