12 дней назад
ML Inference Engineer
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
ML Inference Engineer (LLM/GPU Infrastructure): Building and scaling infrastructure for large-scale LLM workloads with an accent on GPU efficiency, distributed inference, and real-time model serving. Focus on profiling compute, memory, and networking bottlenecks, optimizing latency and throughput, and designing multi-GPU systems across Kubernetes clusters.
Location: San Francisco, United States
Company
Stanford-spun AI startup operating at eight-figure revenue and developing production-scale LLM infrastructure.
What you will do
- Build infrastructure for large-scale LLM workloads and real-time model serving.
- Design distributed inference across single- and multi-GPU systems.
- Optimize latency, throughput, GPU efficiency, scheduling, orchestration, and resource utilization.
- Profile and remove bottlenecks across compute, memory, and networking.
- Scale production workloads across Kubernetes and GPU clusters.
- Make low-level architecture decisions for performance-critical systems.
Requirements
- Strong Python and/or C++ skills.
- Experience with distributed systems or high-performance computing.
- Knowledge of LLM inference, model serving, and modern ML infrastructure.
- Experience optimizing GPU-heavy workloads.
- Exposure to CUDA, NCCL, or Triton and strong understanding of PyTorch.
- Experience with vLLM, TensorRT-LLM, SGLang, or similar technologies; knowledge of quantization, batching, KV caching, and parallelism.
Culture & Benefits
- Work on difficult AI infrastructure problems at production scale.
- Own significant parts of a new inference architecture.
- Work at the intersection of LLMs, GPUs, and distributed systems.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →