Мэтч & Сопровод
Покажет вашу совместимость и напишет письмо
Описание вакансии
Senior Backend Engineer (remote)
August 18, 2026
Senior Backend Engineer (remote)
Company: Claven Claven AI (https://claven.ai/) is a Tashkent based startup building a Sovereign AI Infrastructure & Orchestration Platform — a system that turns bare-metal or managed GPU clusters into cloud-native, efficient, and profitable AI clouds. We make it easy to run inference, fine-tune models, and manage GPU workloads at scale, with a strong focus on data sovereignty for all customers in Uzbekistan and around the globe.
We target enterprise and government customers in regulated industries — banking, finance, healthcare, defense, academia — and already have pilot projects in the pipeline. Our edge is performance, cost efficiency, modern architecture, and a simplified integration experience across both LLM and traditional ML workloads.
We work in a transparent, low-ego, technology-first way, with a culture of hard work, camaraderie, and customer obsession.
How We Work
We are an AI-native team. We use Claude Code daily, alongside other agentic AI tools and workflows to ship faster. This is not a perk, it is how the work happens here.
If you join us, we expect that:
- You actively use Claude Code, Cursor, or similar agentic AI in your daily workflow — not as a curiosity, but as your default way of working.
- When handed an unfamiliar system (GPU scheduling internals, a model-serving engine, a new operator or controller), you go deep on it with AI as your accelerator and come back with answers, prototypes, and opinions.
- You take ownership. You don't wait to be told what to do once a problem is in your lane.
- You maintain curiosity and learning velocity. What You'll Own You will own the backend services and Kubernetes machinery that turn a GPU cluster into a multi-tenant AI platform. This is real distributed-systems and Kubernetes-API work, not CRUD. Specifically:
- Build and evolve our core Go services that orchestrate AI/ML workloads on Kubernetes — routing inference traffic, provisioning and reconciling workloads through the Kubernetes API, and managing their lifecycle.
- Write controllers, operators, admission webhooks, and custom resources (controller-runtime / kubebuilder), including GPU-aware scheduling and placement.
- Integrate and tune the model-serving stack for cost, latency, and throughput, including multi-GPU (tensor-parallel) serving and the model artifact lifecycle.
- Implement multi-tenancy, network policy, and workload isolation in support of our data-sovereignty and air-gapped operating modes — per-tenant isolation, quotas, and RBAC.
- Own the API contracts the platform exposes, and partner with our frontend engineers on the seams where the backend meets the UI, including the gateway and identity layers.
- Help diagnose and resolve production issues across the stack as we onboard pilot customers. What We're Looking For
- We care more about what you can actually do than how many years it took you to get there. The bar is:
- Strong, production-grade Go — you've built and operated real services, not just scripts.
- Real Kubernetes depth: you understand the API and controller/reconciliation model, and you've operated clusters — debugging incidents, not just kubectl apply-ing to a cluster someone else runs.
- Strong Linux, networking, and systems fundamentals.
- Hands-on with infrastructure-as-code, observability tooling, and CI/CD pipelines.
- Comfortable owning an API contract and the boundary it sits on; you write things down and communicate clearly.
- Strong written English and clear, low-ego communication.
- 4+ years in backend, platform, infrastructure, or SRE roles is a useful soft floor, not a hard gate. Bonus Points
- None of these are required. They will accelerate your ramp-up, and we are happy to hire someone with strong fundamentals and zero GPU experience who is excited to learn.
- Writing Kubernetes controllers, operators, admission webhooks, or custom resources (controller-runtime / kubebuilder).
- Hands-on with the NVIDIA stack (CUDA, MIG, NCCL, DCGM) or GPU partitioning and virtualization (MIG, time-slicing, MPS).
- Serving LLMs in production (vLLM, TGI, SGLang, or similar).
- Multi-tenancy, network policy, or compliance work in regulated industries.
- Python — some of our services are Python; comfort crossing over helps.
- API-gateway or identity/OIDC work.
- Background in ML platforms, developer infrastructure, or internal PaaS-style products; familiarity with usage-based metering and quota systems. What You Get
- Ownership of a critical layer of the platform from day one.
- Direct work with the founders and a small, AI-native engineering team — no bureaucracy, no political layers.
- A real customer pipeline in regulated markets — pilots, not vaporware.
- Fully remote, globally. We hire for talent and attitude. Working hours should overlap meaningfully with the team.
- Compensation is discussed during the hiring process and is competitive for the candidate's market. apply
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Текст вакансии взят без изменений