5 дней назад
Inference Infrastructure Architect (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Inference Infrastructure Architect (AI) (vLLM/Kubernetes): Building a bare-metal inference platform for serverless open-weight models and dedicated enterprise deployments across a B300 GPU fleet with an accent on throughput, latency SLOs, cost per token, and multi-tenant serving. Focus on designing Kubernetes-based serving pools, optimizing vLLM and SGLang, implementing KV-cache-aware routing and autoscaling, and integrating observability with production infrastructure.
Location: Remote from mainland China; no relocation required
Company
builds global connectivity infrastructure, including a private multi-cloud IP network, edge technology, and APIs for communications and AI workloads.
What you will do
- Operate and expand a B300 GPU fleet while improving inference throughput, reliability, latency, and cost per token.
- Build serverless serving pools for open-weight models using vLLM and SGLang with continuous batching, prefix caching, low-precision serving, and MoE expert parallelism.
- Develop the fleet layer on Kubernetes Gateway API with KV-cache-aware routing, prefill/decode disaggregation, KV tiering, and multi-LoRA serving.
- Build Kubernetes-on-bare-metal infrastructure covering GPU scheduling, tenant isolation, lifecycle management, model distribution, warm pools, and inference-driven autoscaling.
- Deliver dedicated enterprise deployments with per-tenant pools, GPU-hour metering, latency SLOs, private networking, adapter versioning, canary rollout, and rollback.
- Measure and optimize system performance using DCGM, Prometheus, OpenTelemetry, roofline analysis, batching curves, and cost-per-token metrics.
Requirements
- Production ownership of LLM serving under meaningful traffic and latency constraints; experience at thousands of GPUs or millions of requests per day is valued.
- End-to-end experience operating Kubernetes on GPU fleets, including GPU Operator, device plugins, topology-aware placement, gang scheduling, node pools, and GitOps rollouts.
- Deep operational experience with vLLM or SGLang, including parallelism, quantization, batching, KV-cache configuration, prefix caching, and disaggregation.
- System-level performance engineering skills, including reading engine and DCGM metrics, applying roofline analysis, and sizing deployments quantitatively.
- Python and Go automation experience with strong Linux, networking, and storage knowledge.
- Professional communication and documentation in English and Chinese are required.
Nice to have
- Fine-tuning or reinforcement-learning infrastructure, LoRA pipelines, evaluation harnesses, and champion/challenger rollouts.
- Real-time voice latency optimization and sub-second time-to-first-token targets.
- Multi-region deployments and data-residency experience.
- Contributions to or substantial production use of relevant open-source inference, Kubernetes, networking, or bare-metal projects.
- Participation in CNCF, OpenInfra, vLLM, or SGLang communities.
Culture & Benefits
- Remote work from mainland China with a global, async-friendly team.
- Work with a globally expanding B300 fleet operated from bare metal through customer endpoints.
- Founding role in the China team with influence over platform architecture, hiring, and team practices.
- Open-source contributions are part of the role, with conference travel support.
- Visa sponsorship is available if relocation is later desired through hiring entities in the Netherlands, United States, Ireland, or Saudi Arabia.
Hiring process
- Provide a brief description of an inference system personally improved, including the bottleneck, intervention, and measured result.
- An anonymized example is acceptable.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →