Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Inference Systems Engineer (LLM): Building distributed serving and high-performance networking systems for healthcare-focused large language models with an accent on prefill/decode disaggregation, KV-cache transfer, GPU communication, and production reliability. Focus on optimizing multi-node inference, benchmarking latency and throughput, deploying serving systems on Kubernetes, and diagnosing complex distributed failures.
Location: Menlo Park, California, United States; onsite five days per week
Company
Hippocratic AI is building a safety-focused, healthcare-only large language model platform designed to improve patient outcomes.
What you will do
- Design and operate disaggregated LLM serving architectures with independent prefill and decode pools, KV-cache transfer, routing, batching, and failure recovery.
- Optimize multi-node inference across GPU systems and high-performance networking technologies including NVLink, NVSwitch, InfiniBand, RoCEv2, GPUDirect RDMA, NCCL, and UCX.
- Build and tune serving systems with SGLang, vLLM, NVIDIA Dynamo, TensorRT-LLM, or comparable frameworks.
- Profile compute, memory, network, and storage performance to improve time to first token, inter-token latency, throughput, tail latency, and cost per token.
- Develop benchmarking and capacity-planning workflows across models, GPU types, parallelism strategies, and concurrency levels.
- Productionize Kubernetes-based serving with health checks, autoscaling, safe rollouts, metrics, tracing, diagnostics, and distributed-failure troubleshooting.
Requirements
- Production experience with distributed systems, high-performance computing, or large-scale ML inference platforms.
- Strong understanding of LLM inference, including tensor, pipeline, and data parallelism, continuous batching, KV-cache management, and prefill/decode behavior.
- Hands-on experience with GPU communication and networking technologies such as NCCL, NVLink/NVSwitch, InfiniBand, RoCE, RDMA, or UCX.
- Strong Python skills and working proficiency in C++ or another systems language.
- Ability to profile and debug performance across application, runtime, kernel, network, and infrastructure layers.
- Experience deploying and operating production workloads on Kubernetes with reliable observability standards.
Nice to have
- Experience with disaggregated serving, KV-cache transfer, NVIDIA Dynamo/NIXL, SGLang, vLLM, or TensorRT-LLM.
- Experience with NVLS, SHARP, GPUDirect RDMA, UCX, RDMA congestion control, or GPU-cluster topology optimization.
- CUDA, Triton, custom kernel, or low-level GPU performance experience.
- Experience with speculative decoding, multi-LoRA serving, quantization, or cache-aware routing.
- Contributions to open-source inference, networking, or distributed-systems projects.
Culture & Benefits
- Close collaboration across research, infrastructure, application, product, and model engineering.
- Work on a safety-focused healthcare AI platform.
- Equal opportunity workplace with accommodations available during the hiring process.
- Onsite collaboration in the Menlo Park office five days per week.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
Baseten
4 дня назад
Distributed Systems Engineer (AI Inference)
180 000 - 360 000$
8 дней назад
Forward Deployed Engineer (AI)
200 000 - 400 000$
Nebius
4 дня назад
Senior Software Developer (AI)
4 дня назад
Staff Engineer (AI)
4 дня назад
Software Engineer (AI)
230 000 - 390 000$
12 часов назад
Principal Applied AI Engineer (Healthcare)
230 000 - 280 000$