Мэтч
Покажет вашу совместимость с вакансией
Описание вакансии
Head of Infrastructure
Location: Europe, MENA, LATAM, USA
Salary: Up to 10000 USDT
Company overview
Unimatch Lab is an AI-driven venture studio from Silicon Valley.
We are building our own AI technology stack and a portfolio of AI-driven assets: from consumer AI products and smart devices to on-prem LLM clusters and foundational layers (memory layer, RAM/VRAM optimization, orchestration, AI software for infrastructure and compute).
This includes R&D in local, distributed, and orbital data centers and distributed computing.
Goal: join the world's top 50 AI companies with a combined asset valuation of $10B+ by 2032.
We are an OKR-driven company: measurable outcomes, speed, transparency, and ownership.
We hire A-players: autonomous, fast, strong execution, high accountability.
Position summary
You own the infrastructure that runs self-hosted open-weight LLMs as production services on local GPU clusters. This is a Principal-level Platform / Infrastructure role: you design the system, ship it, operate it, and stay accountable for the trade-offs between latency, throughput, cost, and memory.
You work across on-prem and cloud (AWS and GCP) and treat inference, fine-tuning, GitOps, and reliability as one platform, not a lab experiment. Time goes into cluster and GPU capacity, inference and training pipelines, IaC and multi-cluster Kubernetes, observability, and the security boundary around models, secrets, and supply chain.
The load is high. You make architecture calls with incomplete information and own the outcome. English is the working language, written and spoken.
Key responsibilities
LLM as a production service
- Deploy, operate, and evolve self-hosted open-weight LLMs on local GPU clusters, including inference and fine-tuning pipelines
- Design RAG pipelines and model-serving paths that hold SLO under production traffic
- Choose and justify latency, throughput, cost, and memory trade-offs, and keep those numbers visible in production
GPU and compute platform
- Build and run GPU infrastructure with CUDA-level reasoning: scheduling, utilization, memory, and failure modes
- Plan capacity for high-density compute and keep hybrid paths (on-prem and cloud) operable as one system
Infrastructure, GitOps, and delivery
- Own production-scale Kubernetes (including multi-cluster), Terraform, and ArgoCD as the default change path
- Design GitLab CI and custom pipelines so platform changes are reviewable, repeatable, and reversible
- Keep Linux, networking, and performance work at the depth the cluster actually needs
Reliability, observability, and security
- Run Prometheus, Grafana, and OpenTelemetry (logs, traces, metrics) against SLOs, SLAs, and error budgets
- Lead incident response and postmortems for the LLM and platform surface
- Own IAM, RBAC, zero-trust controls, and secrets (Vault / KMS / SSM), including secure CI/CD and supply-chain security
Required qualifications
- 10+ years in DevOps, Infrastructure, or Platform Engineering, with confirmed work building and running large infrastructure in enterprise or large corporations
- Production experience operating self-hosted open-weight LLMs on local GPU clusters: inference, fine-tuning, and day-2 operations as a live service
- Architectural ownership: you make the call, write it down, and live with the production result
- Comfortable working with high accountability and incomplete information
- Hands-on with Hugging Face, LoRA / QLoRA, and RAG pipelines in production
- Deep Linux (internals, networking, performance), Docker, and production-scale Kubernetes, including multi-cluster
- Terraform as a default; Ansible or Pulumi in real use
- GitLab CI and custom pipelines you have designed and run in production
- Enterprise-level AWS and GCP
- GPU infrastructure with CUDA-level thinking; you can reason about compute and memory under load
- Vector databases in production paths: Qdrant, Pinecone, or Weaviate
- Prometheus, Grafana, OpenTelemetry; SLO / SLA / error budgets; incident response and postmortems
- IAM, RBAC, zero-trust, secrets management (Vault / KMS / SSM), secure CI/CD and supply-chain security
- English B2, written and spoken; Russian, written and spoken
Preferred qualifications
- STEM degree (science, technology, engineering, or mathematics)
- Strong foundation in mathematics, physics, computer science, or engineering
- Bare metal, on-prem, and hybrid infrastructure
- Data center operations and high-density compute
- Distributed computing
- Tenure in Big Tech or large enterprise
- Infrastructure built specifically for AI-first products and AI compute
Technology stack
- Linux, Docker, Kubernetes (production-scale, multi-cluster)
- Terraform (required), Ansible / Pulumi
- GitOps: ArgoCD (required)
- CI/CD: GitLab CI, custom pipelines
- Multi-cloud: AWS and GCP (enterprise-level)
- GPU infrastructure, CUDA, self-hosted LLM (inference and fine-tuning)
- Hugging Face, LoRA / QLoRA, RAG pipelines
- Vector databases: Qdrant / Pinecone / Weaviate
- Prometheus, Grafana, OpenTelemetry
- Vault / KMS / SSM, IAM, RBAC
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Вакансия размещена на Hirify напрямую от HR/нанимающего менеджера