1 час назад
Inference Performance Engineer (LLM Inference)
180 000 - 360 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Inference Performance Engineer (LLM Inference): Building and optimizing production inference systems for demanding AI workloads with an accent on runtime internals, scheduling, routing, and GPU efficiency. Focus on profiling end-to-end latency, implementing quantization and speculative decoding, tuning new model architectures and hardware, and developing benchmarks for cost-effective model serving.
Location: Hybrid in San Francisco, Montreal, New York, Seattle, or Toronto
Salary: $180K–$360K annually, plus equity
Company
Baseten provides inference infrastructure and developer tooling that help AI companies deploy cutting-edge models in production.
What you will do
- Implement and productionize inference techniques including quantization, speculative decoding, KV-cache reuse, chunked prefill, LoRA, and guided generation.
- Profile and optimize inference across runtime internals, GPU kernels, memory layout, scheduling, batching, routing, and prefill/decode disaggregation.
- Improve tokens per GPU-hour, utilization, latency, throughput, and serving costs.
- Bring up and tune new model architectures on emerging hardware.
- Build benchmarking frameworks across model architectures, batch sizes, sequence lengths, and hardware configurations.
- Contribute to open-source inference engines and collaborate with model, infrastructure, and customer-facing teams.
Requirements
- Bachelor's, Master's, or Ph.D. in Computer Science, Engineering, Mathematics, or a related field.
- Experience with Python, C++, or another general-purpose programming language.
- Familiarity with LLM optimization techniques such as quantization, speculative decoding, and continuous batching.
- Strong familiarity with ML libraries, especially PyTorch, TensorRT, or TensorRT-LLM.
- Demonstrated interest and experience in LLMs.
- Deep understanding of GPU architecture.
Nice to have
- Contributions to vLLM, SGLang, TensorRT-LLM, or another inference engine.
- Experience with large-scale distributed serving, autoscaling, load balancing, or multi-region and multi-cloud capacity.
- Experience optimizing GPU kernels with CUDA, Triton, CUTLASS, or similar technologies.
- Production experience with FP8/FP4, AWQ, GPTQ, or speculative decoding.
- Experience developing and deploying AI/ML inference solutions.
Culture & Benefits
- Competitive compensation with meaningful equity.
- Flexible PTO and a company-wide winter break.
- Paid parental leave.
- Fertility and family-building stipend through Carrot.
- U.S. employees receive company-facilitated 401(k) and full medical, dental, and vision coverage for employees and dependents.
- Exposure to a range of ML startups and opportunities for learning and networking.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
6 дней назад
Machine Learning Engineer (AI Inference)
275 000 - 300 000$
Microsoft AI
6 дней назад
Member of Technical Staff, Inference Systems Research (AI)
142 800 - 274 800$
6 дней назад
Forward Deployed Engineer (Inference)
180 000 - 300 000$
5 дней назад
ML Performance Engineer (AI)
100 000 - 150 000$
3 дня назад
Senior Research Engineer (AI)
6 дней назад
Post-Training Engineer (AI)
300 000 - 350 000$