2 дня назад
ML Systems Engineer (Inference Infrastructure)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
ML Systems Engineer (Inference Infrastructure): Own and optimize the cost and performance of the inference stack, focusing on caching, batching, quantization, decoding, and kernel-level optimization. Focus on improving throughput, latency, and reliability while working with serving engines like vLLM, SGLang, and TensorRT-LLM.
Location
Location: San Francisco, hybrid work format
Company
is a startup focused on building efficient, adaptable AI systems that evolve in real-time to expand access and innovation.
What you will do
- Own cost and performance of the inference stack, optimizing caching, batching, quantization, decoding, and kernel-level operations.
- Improve throughput, cost, and tail latency through KV-cache management, continuous batching, speculative decoding, and quantization.
- Optimize long-context prefill and decode workloads based on real production traffic.
- Tune routing between infrastructure and external providers based on cost, capacity, and performance.
- Work within serving engines such as vLLM, SGLang, and TensorRT-LLM, including low-level framework optimizations.
- Build profiling and measurement systems to analyze time, memory, and compute usage.
Requirements
- 5+ years experience in ML systems, inference infrastructure, or performance engineering with measurable improvements in cost or latency.
- Deep understanding of model serving, including prefill, decode, memory bandwidth, batching, and concurrency.
- Production experience with serving engines such as vLLM, SGLang, or TensorRT-LLM.
- Strong Python skills and proficiency in C++, Rust, or another systems language.
- Experience with GPU performance including CUDA, NCCL, mixed precision, memory layout, kernels, or quantization.
Culture & Benefits
- Flexible work with in-person collaboration in the Bay Area and a distributed global-first team.
- Annual travel stipend called Passport to explore new countries.
- Weekly lunch stipend for take-out or grocery delivery.
- Comprehensive medical benefits and generous paid time off.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
3 дня назад
ML Engineer (AI)
3 дня назад
Generalist Engineer (AI)
3 дня назад
Performance Engineer (AI)
200 000 - 400 000$
3 дня назад
ML Engineer (AI Robotics)
200 000 - 350 000$
3 дня назад
Software Engineer (AI Inference & RL Systems)
225 000 - 550 000$
3 дня назад
Compiler Engineer (AI Infrastructure)
180 000 - 400 000$