5 дней назад
Senior Machine Learning Engineer (LLM Inference Optimization)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Machine Learning Engineer (LLM Inference Optimization): Building fast, reliable, and cost-efficient inference services for frontier models with an accent on model optimization, serving architecture, benchmarking, and production deployment. Focus on optimizing latency, throughput, memory efficiency, GPU utilization, model quality, and cost per token across inference engines, compression workflows, and distributed serving systems.
Location: London, United Kingdom
Company
Nebius is building a full-stack AI cloud platform for model training, inference, and production deployment, with infrastructure spanning compute, storage, networking, and applied AI.
What you will do
- Own optimization work for model families, customer endpoints, and serving backends.
- Deploy, configure, benchmark, and extend inference engines including vLLM, SGLang, TensorRT-LLM, Triton Inference Server, and NVIDIA Dynamo.
- Optimize LLM and VLM endpoints for latency, throughput, memory efficiency, GPU utilization, quality, reliability, and cost per token.
- Build and productionize model-compression workflows covering quantization, quantization-aware training, distillation, low-bit serving, and accuracy recovery.
- Implement inference acceleration and serving techniques such as speculative decoding, KV-cache optimization, prefix caching, chunked prefill, continuous batching, and disaggregated prefill/decode.
- Build reproducible benchmarks, diagnose bottlenecks across model, kernel, runtime, scheduler, gateway, and cluster layers, and document rollout plans and performance results.
Requirements
- Strong Python and PyTorch engineering skills.
- Hands-on experience deploying or optimizing LLM, VLM, or high-throughput transformer inference systems.
- Practical knowledge of at least one modern inference stack, such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, or KServe.
- Strong understanding of transformer inference bottlenecks, including KV cache, attention, memory bandwidth, batching, parallelism, and long-context serving.
- Ability to reason quantitatively about latency, throughput, quality, utilization, and cost trade-offs.
- Applicants must be authorized to work in the United Kingdom and provide proof of employment eligibility.
Nice to have
- Experience with quantization, FP8, INT8, INT4, NVFP4, MXFP4, AWQ, GPTQ, SmoothQuant, distillation, speculative decoding, EAGLE, Medusa, or multi-token prediction.
- Experience with agentic workloads, tool calling, structured outputs, streaming APIs, high concurrency, and multi-step orchestration.
- CUDA or Triton familiarity.
- Open-source contributions to inference and ML infrastructure projects such as vLLM, SGLang, TensorRT-LLM, FlashInfer, PyTorch, Triton, Ray, or KServe.
Culture & Benefits
- Competitive compensation and career growth opportunities.
- Flexibility, ownership, and opportunities for learning.
- Collaborative, innovative, and international working environment.
- Opportunity to work on impactful AI infrastructure projects.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
6 дней назад
LLM Inference Engineer
6 дней назад
Senior Machine Learning Engineer (AI)
12 дней назад
Machine Learning Engineer (LLM Inference Serving, vLLM)
150 000 - 190 000$
12 дней назад
Member of Technical Staff, Inference Performance (AI)
200 000$
7 дней назад
LLM Engineer (AI)
72 000 - 100 000$
12 дней назад