Назад
5 дней назад

Senior Machine Learning Engineer (LLM Inference Optimization)

Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
UK/US/Netherlands +2 еще
Вакансия из списка Hirify.GlobalВакансия из Hirify RU Global, списка компаний с восточно-европейскими корнями
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Machine Learning Engineer (LLM Inference Optimization): Building fast, reliable, and cost-efficient inference services for frontier models with an accent on model optimization, serving architecture, benchmarking, and production deployment. Focus on optimizing latency, throughput, memory efficiency, GPU utilization, model quality, and cost per token across inference engines, compression workflows, and distributed serving systems.

Location: London, United Kingdom

Company

Nebius is building a full-stack AI cloud platform for model training, inference, and production deployment, with infrastructure spanning compute, storage, networking, and applied AI.

What you will do

  • Own optimization work for model families, customer endpoints, and serving backends.
  • Deploy, configure, benchmark, and extend inference engines including vLLM, SGLang, TensorRT-LLM, Triton Inference Server, and NVIDIA Dynamo.
  • Optimize LLM and VLM endpoints for latency, throughput, memory efficiency, GPU utilization, quality, reliability, and cost per token.
  • Build and productionize model-compression workflows covering quantization, quantization-aware training, distillation, low-bit serving, and accuracy recovery.
  • Implement inference acceleration and serving techniques such as speculative decoding, KV-cache optimization, prefix caching, chunked prefill, continuous batching, and disaggregated prefill/decode.
  • Build reproducible benchmarks, diagnose bottlenecks across model, kernel, runtime, scheduler, gateway, and cluster layers, and document rollout plans and performance results.

Requirements

  • Strong Python and PyTorch engineering skills.
  • Hands-on experience deploying or optimizing LLM, VLM, or high-throughput transformer inference systems.
  • Practical knowledge of at least one modern inference stack, such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, or KServe.
  • Strong understanding of transformer inference bottlenecks, including KV cache, attention, memory bandwidth, batching, parallelism, and long-context serving.
  • Ability to reason quantitatively about latency, throughput, quality, utilization, and cost trade-offs.
  • Applicants must be authorized to work in the United Kingdom and provide proof of employment eligibility.

Nice to have

  • Experience with quantization, FP8, INT8, INT4, NVFP4, MXFP4, AWQ, GPTQ, SmoothQuant, distillation, speculative decoding, EAGLE, Medusa, or multi-token prediction.
  • Experience with agentic workloads, tool calling, structured outputs, streaming APIs, high concurrency, and multi-step orchestration.
  • CUDA or Triton familiarity.
  • Open-source contributions to inference and ML infrastructure projects such as vLLM, SGLang, TensorRT-LLM, FlashInfer, PyTorch, Triton, Ray, or KServe.

Culture & Benefits

  • Competitive compensation and career growth opportunities.
  • Flexibility, ownership, and opportunities for learning.
  • Collaborative, innovative, and international working environment.
  • Opportunity to work on impactful AI infrastructure projects.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →