Назад
2 дня назад

Senior Machine Learning Engineer (LLM Inference Optimization)

195 200 - 262 200$
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify RU Global, списка компаний с восточно-европейскими корнями
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/

TL;DR

Senior Machine Learning Engineer (LLM Inference Optimization): Building and optimizing high-performance LLM and VLM endpoints for a full-stack AI cloud platform with an accent on latency, throughput, and memory efficiency. Focus on productionizing model-compression workflows, implementing advanced inference techniques like speculative decoding, and diagnosing GPU bottlenecks.

Location: Palo Alto, California, United States. Applicants must be authorized to work in the USA.

Salary: $195,200 - $262,200 USD

Company

Nebius is building a full-stack AI cloud platform designed to support developers and enterprises from model training through to production deployment.

What you will do

  • Own optimization projects for specific model families, customer endpoints, and serving backends.
  • Optimize LLM and VLM endpoints for latency, throughput, memory efficiency, and cost per token.
  • Deploy, configure, and extend inference engines such as vLLM, SGLang, TensorRT-LLM, and Triton Inference Server.
  • Productionize model-compression workflows, including quantization, distillation, and low-bit serving.
  • Implement advanced techniques like speculative decoding, KV-cache optimization, and continuous batching.
  • Partner with GPU kernel and platform engineers to diagnose bottlenecks across the system stack.

Requirements

  • Strong Python and PyTorch engineering skills.
  • Hands-on experience deploying or optimizing high-throughput transformer inference systems.
  • Practical knowledge of modern inference stacks (vLLM, SGLang, TensorRT-LLM, Triton, Ray Serve, etc.).
  • Deep understanding of transformer bottlenecks, including KV cache, attention, and memory bandwidth.
  • Ability to reason quantitatively about latency, throughput, and quality tradeoffs.
  • Authorization to work in the United States is required.

Nice to have

  • Experience with advanced quantization techniques (FP8, INT8, INT4, AWQ, GPTQ, SmoothQuant).
  • Experience with distillation, speculative decoding (EAGLE, Medusa), or multi-token prediction.
  • Familiarity with CUDA or Triton kernels.
  • Open-source contributions to projects like vLLM, SGLang, PyTorch, or Ray.

Culture & Benefits

  • 100% company-paid medical, dental, and vision coverage for employees and families.
  • 401(k) plan with up to 4% company match and immediate vesting.
  • Generous parental leave (20 weeks for primary, 12 weeks for secondary caregivers).
  • Remote work reimbursement for mobile and internet expenses.
  • Collaborative and innovative environment with a focus on ownership and impactful AI projects.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →