Назад
3 дня назад

Inference Systems Engineer (LLM)

Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Inference Systems Engineer (LLM): Building distributed serving and high-performance networking systems for healthcare-focused large language models with an accent on prefill/decode disaggregation, KV-cache transfer, GPU communication, and production reliability. Focus on optimizing multi-node inference, benchmarking latency and throughput, deploying serving systems on Kubernetes, and diagnosing complex distributed failures.

Location: Menlo Park, California, United States; onsite five days per week

Company

Hippocratic AI is building a safety-focused, healthcare-only large language model platform designed to improve patient outcomes.

What you will do

  • Design and operate disaggregated LLM serving architectures with independent prefill and decode pools, KV-cache transfer, routing, batching, and failure recovery.
  • Optimize multi-node inference across GPU systems and high-performance networking technologies including NVLink, NVSwitch, InfiniBand, RoCEv2, GPUDirect RDMA, NCCL, and UCX.
  • Build and tune serving systems with SGLang, vLLM, NVIDIA Dynamo, TensorRT-LLM, or comparable frameworks.
  • Profile compute, memory, network, and storage performance to improve time to first token, inter-token latency, throughput, tail latency, and cost per token.
  • Develop benchmarking and capacity-planning workflows across models, GPU types, parallelism strategies, and concurrency levels.
  • Productionize Kubernetes-based serving with health checks, autoscaling, safe rollouts, metrics, tracing, diagnostics, and distributed-failure troubleshooting.

Requirements

  • Production experience with distributed systems, high-performance computing, or large-scale ML inference platforms.
  • Strong understanding of LLM inference, including tensor, pipeline, and data parallelism, continuous batching, KV-cache management, and prefill/decode behavior.
  • Hands-on experience with GPU communication and networking technologies such as NCCL, NVLink/NVSwitch, InfiniBand, RoCE, RDMA, or UCX.
  • Strong Python skills and working proficiency in C++ or another systems language.
  • Ability to profile and debug performance across application, runtime, kernel, network, and infrastructure layers.
  • Experience deploying and operating production workloads on Kubernetes with reliable observability standards.

Nice to have

  • Experience with disaggregated serving, KV-cache transfer, NVIDIA Dynamo/NIXL, SGLang, vLLM, or TensorRT-LLM.
  • Experience with NVLS, SHARP, GPUDirect RDMA, UCX, RDMA congestion control, or GPU-cluster topology optimization.
  • CUDA, Triton, custom kernel, or low-level GPU performance experience.
  • Experience with speculative decoding, multi-LoRA serving, quantization, or cache-aware routing.
  • Contributions to open-source inference, networking, or distributed-systems projects.

Culture & Benefits

  • Close collaboration across research, infrastructure, application, product, and model engineering.
  • Work on a safety-focused healthcare AI platform.
  • Equal opportunity workplace with accommodations available during the hiring process.
  • Onsite collaboration in the Menlo Park office five days per week.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →