Назад
Company hidden
5 дней назад

Inference Infrastructure Architect (AI)

Формат работы
remote (только China)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US/Ireland/Netherlands +2 еще
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Inference Infrastructure Architect (AI) (vLLM/Kubernetes): Building a bare-metal inference platform for serverless open-weight models and dedicated enterprise deployments across a B300 GPU fleet with an accent on throughput, latency SLOs, cost per token, and multi-tenant serving. Focus on designing Kubernetes-based serving pools, optimizing vLLM and SGLang, implementing KV-cache-aware routing and autoscaling, and integrating observability with production infrastructure.

Location: Remote from mainland China; no relocation required

Company

hirify.global builds global connectivity infrastructure, including a private multi-cloud IP network, edge technology, and APIs for communications and AI workloads.

What you will do

  • Operate and expand a B300 GPU fleet while improving inference throughput, reliability, latency, and cost per token.
  • Build serverless serving pools for open-weight models using vLLM and SGLang with continuous batching, prefix caching, low-precision serving, and MoE expert parallelism.
  • Develop the fleet layer on Kubernetes Gateway API with KV-cache-aware routing, prefill/decode disaggregation, KV tiering, and multi-LoRA serving.
  • Build Kubernetes-on-bare-metal infrastructure covering GPU scheduling, tenant isolation, lifecycle management, model distribution, warm pools, and inference-driven autoscaling.
  • Deliver dedicated enterprise deployments with per-tenant pools, GPU-hour metering, latency SLOs, private networking, adapter versioning, canary rollout, and rollback.
  • Measure and optimize system performance using DCGM, Prometheus, OpenTelemetry, roofline analysis, batching curves, and cost-per-token metrics.

Requirements

  • Production ownership of LLM serving under meaningful traffic and latency constraints; experience at thousands of GPUs or millions of requests per day is valued.
  • End-to-end experience operating Kubernetes on GPU fleets, including GPU Operator, device plugins, topology-aware placement, gang scheduling, node pools, and GitOps rollouts.
  • Deep operational experience with vLLM or SGLang, including parallelism, quantization, batching, KV-cache configuration, prefix caching, and disaggregation.
  • System-level performance engineering skills, including reading engine and DCGM metrics, applying roofline analysis, and sizing deployments quantitatively.
  • Python and Go automation experience with strong Linux, networking, and storage knowledge.
  • Professional communication and documentation in English and Chinese are required.

Nice to have

  • Fine-tuning or reinforcement-learning infrastructure, LoRA pipelines, evaluation harnesses, and champion/challenger rollouts.
  • Real-time voice latency optimization and sub-second time-to-first-token targets.
  • Multi-region deployments and data-residency experience.
  • Contributions to or substantial production use of relevant open-source inference, Kubernetes, networking, or bare-metal projects.
  • Participation in CNCF, OpenInfra, vLLM, or SGLang communities.

Culture & Benefits

  • Remote work from mainland China with a global, async-friendly team.
  • Work with a globally expanding B300 fleet operated from bare metal through customer endpoints.
  • Founding role in the China team with influence over platform architecture, hiring, and team practices.
  • Open-source contributions are part of the role, with conference travel support.
  • Visa sponsorship is available if relocation is later desired through hiring entities in the Netherlands, United States, Ireland, or Saudi Arabia.

Hiring process

  • Provide a brief description of an inference system personally improved, including the bottleneck, intervention, and measured result.
  • An anonymized example is acceptable.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →