Назад
Company hidden
12 дней назад

Machine Learning Engineer (LLM Inference Serving, vLLM)

150 000 - 190 000$
Формат работы
onsite
Тип работы
fulltime
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Machine Learning Engineer (LLM Inference Serving, vLLM): Refactoring inference framework internals and building distributed serving systems for LLMs with an accent on vLLM/SGLang, throughput and latency optimization, and KV cache lifecycle management. Focus on inference orchestration with Ray or Dynamo, scheduling, and improving production inference performance.

Location: Santa Clara, California, United States — onsite

Salary: $150,000–$190,000 per year

Company

Stealth AI infrastructure startup developing systems for large language model inference and serving.

What you will do

  • Build and optimize LLM inference serving systems.
  • Refactor inference framework internals to improve performance.
  • Manage the KV cache lifecycle.
  • Optimize distributed serving throughput and latency.
  • Work on scheduling and inference orchestration using Ray, Dynamo, or similar platforms.

Requirements

  • Deep hands-on experience with vLLM or SGLang.
  • Public GitHub contributions to vLLM or SGLang.
  • Experience improving inference framework performance.
  • Experience with KV cache lifecycle management.
  • Experience with Ray, Dynamo, or similar inference orchestration platforms.
  • Ability to work onsite in Santa Clara, California.

Culture & Benefits

  • Work on AI infrastructure for LLM inference.
  • Focus on distributed serving, scheduling, throughput, and latency optimization.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →