8 дней назад
Senior Backend Engineer – LLM Inference
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Backend Engineer – LLM Inference (Go/Kubernetes): Building reliable backend services and Kubernetes-native applications for LLM inference on an AI cloud with an accent on distributed systems, API design, scalability, and production operations. Focus on designing concurrency-safe services, implementing reconciliation loops and reliable event processing, and improving observability and failure handling across the platform.
Location: Based in Helsinki, Finland or London, UK, or remote in Europe; hybrid work mode
Company
is building a full-stack AI cloud spanning data centers, hardware, and cloud infrastructure for AI teams.
What you will do
- Design, build, and maintain production backend services in Go.
- Develop Kubernetes-native applications and automation for reliable platform operations.
- Build APIs and integrations supporting LLM inference services and the wider cloud platform.
- Improve scalability, performance, reliability, testing, observability, and operational processes.
- Collaborate with infrastructure and AI engineering teams on technical design, delivery, and production troubleshooting.
Requirements
- Senior-level experience designing, shipping, and operating production backend services.
- Strong production experience with Go, including concurrency, context propagation, API design, and error handling.
- Knowledge of asynchronous execution, HTTP, long-lived connections, distributed systems, caching, consistency, retries, idempotency, and event processing.
- Strong SQL and relational data modelling skills, including transactions, schema migrations, and PostgreSQL.
- Hands-on Kubernetes-native development and deployment experience, including controllers, operators, custom resources, reconciliation loops, Helm, networking, readiness, configuration, and safe rollouts.
- Clear written English and the ability to explain design decisions, trade-offs, and operational behaviour.
Nice to have
- Rust experience, especially in asynchronous or network services.
- Experience with API gateways, networking, traffic management, Valkey/Redis, NATS JetStream, or comparable systems.
- Familiarity with LLM inference systems and protocols such as vLLM, SGLang, NVIDIA Dynamo, or llm-d.
- Experience with OpenTelemetry, Prometheus, capacity planning, service discovery, GitOps, multi-cluster deployments, or Python integration tooling.
Culture & Benefits
- Full-time, permanent employment.
- Cash and equity compensation with healthcare, lunch, wellbeing, and other benefits.
- Work alongside engineers, researchers, and partners across the AI ecosystem.
- International environment with more than 40 nationalities.
Hiring process
- Applications are submitted through the Careers page; email applications are not accepted.
- Hiring continues until the position is filled.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →