Назад
обновлено 56 минут назад

Head of Infrastructure

10 000USDT
Формат работы
remote (Global)
Тип работы
fulltime
Грейд
lead
Английский
b2
Страна
US/Europe
Вакансия от Hirify. Размещена напрямую Вакансия размещена на Hirify напрямую от HR/нанимающего менеджера

Мэтч

Покажет вашу совместимость с вакансией

Описание вакансии

TL;DR
Head of Infrastructure (AI): Designing and operating production-scale infrastructure for self-hosted LLMs on local GPU clusters with an accent on high-density compute and reliability. Focus on building robust inference pipelines, managing multi-cluster Kubernetes environments, and ensuring security across the AI supply chain.

Head of Infrastructure

Location: Europe, MENA, LATAM, USA

Salary: Up to 10000 USDT

Company overview

Unimatch Lab is an AI-driven venture studio from Silicon Valley.

We are building our own AI technology stack and a portfolio of AI-driven assets: from consumer AI products and smart devices to on-prem LLM clusters and foundational layers (memory layer, RAM/VRAM optimization, orchestration, AI software for infrastructure and compute).

This includes R&D in local, distributed, and orbital data centers and distributed computing.

Goal: join the world's top 50 AI companies with a combined asset valuation of $10B+ by 2032.

We are an OKR-driven company: measurable outcomes, speed, transparency, and ownership.

We hire A-players: autonomous, fast, strong execution, high accountability.

Position summary

You own the infrastructure that runs self-hosted open-weight LLMs as production services on local GPU clusters. This is a Principal-level Platform / Infrastructure role: you design the system, ship it, operate it, and stay accountable for the trade-offs between latency, throughput, cost, and memory.

You work across on-prem and cloud (AWS and GCP) and treat inference, fine-tuning, GitOps, and reliability as one platform, not a lab experiment. Time goes into cluster and GPU capacity, inference and training pipelines, IaC and multi-cluster Kubernetes, observability, and the security boundary around models, secrets, and supply chain.

The load is high. You make architecture calls with incomplete information and own the outcome. English is the working language, written and spoken.

Key responsibilities

LLM as a production service

  • Deploy, operate, and evolve self-hosted open-weight LLMs on local GPU clusters, including inference and fine-tuning pipelines
  • Design RAG pipelines and model-serving paths that hold SLO under production traffic
  • Choose and justify latency, throughput, cost, and memory trade-offs, and keep those numbers visible in production

GPU and compute platform

  • Build and run GPU infrastructure with CUDA-level reasoning: scheduling, utilization, memory, and failure modes
  • Plan capacity for high-density compute and keep hybrid paths (on-prem and cloud) operable as one system

Infrastructure, GitOps, and delivery

  • Own production-scale Kubernetes (including multi-cluster), Terraform, and ArgoCD as the default change path
  • Design GitLab CI and custom pipelines so platform changes are reviewable, repeatable, and reversible
  • Keep Linux, networking, and performance work at the depth the cluster actually needs

Reliability, observability, and security

  • Run Prometheus, Grafana, and OpenTelemetry (logs, traces, metrics) against SLOs, SLAs, and error budgets
  • Lead incident response and postmortems for the LLM and platform surface
  • Own IAM, RBAC, zero-trust controls, and secrets (Vault / KMS / SSM), including secure CI/CD and supply-chain security

Required qualifications

  • 10+ years in DevOps, Infrastructure, or Platform Engineering, with confirmed work building and running large infrastructure in enterprise or large corporations
  • Production experience operating self-hosted open-weight LLMs on local GPU clusters: inference, fine-tuning, and day-2 operations as a live service
  • Architectural ownership: you make the call, write it down, and live with the production result
  • Comfortable working with high accountability and incomplete information
  • Hands-on with Hugging Face, LoRA / QLoRA, and RAG pipelines in production
  • Deep Linux (internals, networking, performance), Docker, and production-scale Kubernetes, including multi-cluster
  • Terraform as a default; Ansible or Pulumi in real use
  • GitLab CI and custom pipelines you have designed and run in production
  • Enterprise-level AWS and GCP
  • GPU infrastructure with CUDA-level thinking; you can reason about compute and memory under load
  • Vector databases in production paths: Qdrant, Pinecone, or Weaviate
  • Prometheus, Grafana, OpenTelemetry; SLO / SLA / error budgets; incident response and postmortems
  • IAM, RBAC, zero-trust, secrets management (Vault / KMS / SSM), secure CI/CD and supply-chain security
  • English B2, written and spoken; Russian, written and spoken

Preferred qualifications

  • STEM degree (science, technology, engineering, or mathematics)
  • Strong foundation in mathematics, physics, computer science, or engineering
  • Bare metal, on-prem, and hybrid infrastructure
  • Data center operations and high-density compute
  • Distributed computing
  • Tenure in Big Tech or large enterprise
  • Infrastructure built specifically for AI-first products and AI compute

Technology stack

  • Linux, Docker, Kubernetes (production-scale, multi-cluster)
  • Terraform (required), Ansible / Pulumi
  • GitOps: ArgoCD (required)
  • CI/CD: GitLab CI, custom pipelines
  • Multi-cloud: AWS and GCP (enterprise-level)
  • GPU infrastructure, CUDA, self-hosted LLM (inference and fine-tuning)
  • Hugging Face, LoRA / QLoRA, RAG pipelines
  • Vector databases: Qdrant / Pinecone / Weaviate
  • Prometheus, Grafana, OpenTelemetry
  • Vault / KMS / SSM, IAM, RBAC

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →

Вакансия размещена на Hirify напрямую от HR/нанимающего менеджера