Назад
2 дня назад

Senior Applied Scientist (Efficient LLM Inference)

Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
UK/Poland/CR +4 еще
Вакансия из списка Hirify.GlobalВакансия из Hirify RU Global, списка компаний с восточно-европейскими корнями
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Applied Scientist (Efficient LLM Inference): Developing and productionizing efficient LLM and VLM inference methods with an accent on quantization, distillation, speculative decoding, KV-cache optimization, and model/runtime co-optimization. Focus on designing rigorous experiments, building PyTorch and Triton prototypes, measuring quality and system performance, and transferring research into production inference components.

Location: Amsterdam, Netherlands; Berlin, Germany; London, United Kingdom; Poland; Prague, Czech Republic; Tel Aviv, Israel; or Zurich, Switzerland. Applicants must be authorized to work in the country where they apply.

Company

Nebius builds a full-stack AI cloud platform for data processing, model training, and production deployment, with infrastructure spanning compute, storage, networking, and applied AI.

What you will do

  • Own research projects from hypothesis and experiment design through ablation, prototyping, and production handoff.
  • Define and execute research programs in efficient LLM and VLM inference with measurable production impact.
  • Develop and productionize methods for quantization, QAT, distillation, speculative decoding, KV-cache reuse and compression, long-context inference, MoE routing, and model/runtime co-optimization.
  • Build prototypes with PyTorch, Triton, CUDA-adjacent tooling, and inference-serving frameworks.
  • Evaluate quality, latency, throughput, numerical stability, memory footprint, tail latency, and cost per token.
  • Collaborate with MLE, GPU kernel, backend infrastructure, product, and customer teams; publish technical work and mentor engineers and scientists.

Requirements

  • PhD in computer science, machine learning, ML systems, computer systems, computer architecture, electrical engineering, applied mathematics, or a related field.
  • Strong publication record or equivalent research artifacts in ML, ML systems, efficient inference, model compression, quantization, distillation, or serving systems.
  • Strong hands-on coding skills in Python and PyTorch.
  • Deep understanding of LLMs, VLMs, transformer inference, decoding algorithms, model compression, quantization, and production-serving tradeoffs.
  • Strong experimental design skills covering ablations, baselines, metrics, statistical reasoning, and failure analysis.
  • Excellent written and verbal communication.

Nice to have

  • First-author publications in leading ML, systems, architecture, or NLP venues.
  • Experience deploying ML models or inference optimizations in production.
  • Experience with vLLM, SGLang, TensorRT-LLM, NVIDIA Dynamo, FlashAttention, FlashInfer, Triton, CUDA, or PyTorch internals.
  • Experience with post-training, SFT, DPO, RLHF, RLAIF, preference optimization, or synthetic data generation related to inference quality or efficiency.
  • Open-source research artifacts, widely used benchmarks, technical blogs, or invited talks in efficient AI systems.

Culture & Benefits

  • Competitive compensation.
  • Career growth and learning opportunities.
  • Flexibility and ownership.
  • Collaborative, innovative, and international working environment.
  • Opportunity to work on impactful AI projects.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →