Назад
4 дня назад

Senior Machine Learning Engineer, LLM Inference Optimization

Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
UK/US/Netherlands +2 еще
Вакансия из списка Hirify.GlobalВакансия из Hirify RU Global, списка компаний с восточно-европейскими корнями
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Machine Learning Engineer, LLM Inference Optimization (AI/LLM Inference): Building fast, reliable, and cost-efficient inference services for frontier models with an accent on latency, throughput, memory efficiency, GPU utilization, and model quality. Focus on optimizing serving engines, implementing compression and decoding techniques, building reproducible benchmarks, and diagnosing bottlenecks across model, runtime, scheduler, gateway, and cluster layers.

Location: Zurich, Switzerland. Applicants must be authorized to work in the country in which they apply and provide proof of employment eligibility.

Company

Nebius builds a full-stack AI cloud platform covering data and model training through production deployment, with expertise in GPU orchestration, inference optimization, compute, storage, networking, and applied AI.

What you will do

  • Own optimization projects for model families, customer endpoints, and serving backends.
  • Deploy, configure, benchmark, and extend inference engines such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, and NVIDIA Dynamo.
  • Optimize LLM and VLM endpoints for latency, throughput, memory efficiency, GPU utilization, quality, and cost per token.
  • Build production-ready model-compression workflows, including quantization, quantization-aware training, distillation, low-bit serving, and accuracy recovery.
  • Implement inference acceleration techniques such as speculative decoding, KV-cache optimization, prefix caching, chunked prefill, continuous batching, and disaggregated prefill/decode serving.
  • Develop reproducible benchmarks and collaborate with kernel and platform engineers to diagnose bottlenecks across the serving stack.

Requirements

  • Strong Python and PyTorch engineering skills.
  • Hands-on experience deploying or optimizing LLM, VLM, or high-throughput transformer inference systems.
  • Practical knowledge of at least one modern inference stack, such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, or KServe.
  • Strong understanding of transformer inference bottlenecks, including KV cache, attention, memory bandwidth, batching, parallelism, and long-context serving.
  • Ability to reason quantitatively about latency, throughput, quality, utilization, and cost trade-offs.
  • Strong communication and collaboration skills across research, kernel, infrastructure, product, and customer teams.

Nice to have

  • Experience with quantization techniques and formats such as FP8, INT8, INT4, NVFP4, MXFP4, AWQ, GPTQ, or SmoothQuant.
  • Experience with distillation, speculative decoding, EAGLE, Medusa, multi-token prediction, or other inference acceleration methods.
  • Experience with agentic workloads, tool calling, structured outputs, streaming APIs, high concurrency, and multi-step orchestration.
  • CUDA or Triton familiarity.
  • Open-source contributions to relevant inference, serving, or machine learning projects.

Culture & Benefits

  • Competitive compensation.
  • Career growth and learning opportunities.
  • Flexibility, ownership, and the opportunity to shape AI infrastructure.
  • Collaborative, innovative, and international environment.
  • Work on impactful AI projects with experienced engineering teams.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →