Senior Machine Learning Engineer (LLM Inference Optimization)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
TL;DR
Senior Machine Learning Engineer (LLM Inference Optimization): Building and optimizing high-performance LLM and VLM endpoints for a full-stack AI cloud platform with an accent on latency, throughput, and memory efficiency. Focus on productionizing model-compression workflows, implementing advanced inference techniques like speculative decoding, and diagnosing GPU bottlenecks.
Location: Palo Alto, California, United States. Applicants must be authorized to work in the USA.
Salary: $195,200 - $262,200 USD
Company
Nebius is building a full-stack AI cloud platform designed to support developers and enterprises from model training through to production deployment.
What you will do
- Own optimization projects for specific model families, customer endpoints, and serving backends.
- Optimize LLM and VLM endpoints for latency, throughput, memory efficiency, and cost per token.
- Deploy, configure, and extend inference engines such as vLLM, SGLang, TensorRT-LLM, and Triton Inference Server.
- Productionize model-compression workflows, including quantization, distillation, and low-bit serving.
- Implement advanced techniques like speculative decoding, KV-cache optimization, and continuous batching.
- Partner with GPU kernel and platform engineers to diagnose bottlenecks across the system stack.
Requirements
- Strong Python and PyTorch engineering skills.
- Hands-on experience deploying or optimizing high-throughput transformer inference systems.
- Practical knowledge of modern inference stacks (vLLM, SGLang, TensorRT-LLM, Triton, Ray Serve, etc.).
- Deep understanding of transformer bottlenecks, including KV cache, attention, and memory bandwidth.
- Ability to reason quantitatively about latency, throughput, and quality tradeoffs.
- Authorization to work in the United States is required.
Nice to have
- Experience with advanced quantization techniques (FP8, INT8, INT4, AWQ, GPTQ, SmoothQuant).
- Experience with distillation, speculative decoding (EAGLE, Medusa), or multi-token prediction.
- Familiarity with CUDA or Triton kernels.
- Open-source contributions to projects like vLLM, SGLang, PyTorch, or Ray.
Culture & Benefits
- 100% company-paid medical, dental, and vision coverage for employees and families.
- 401(k) plan with up to 4% company match and immediate vesting.
- Generous parental leave (20 weeks for primary, 12 weeks for secondary caregivers).
- Remote work reimbursement for mobile and internet expenses.
- Collaborative and innovative environment with a focus on ownership and impactful AI projects.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →