Назад
3 дня назад

Senior ML Systems Engineer Inference (LLM)

150 000 - 220 000$
Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Страна
US
vacancy_detail.hirify_telegram_tooltipВакансия из Telegram канала -

Мэтч & Сопровод

Покажет вашу совместимость и напишет письмо

Описание вакансии

TL;DR
Senior ML Systems Engineer Inference (LLM): Leading end-to-end LLM serving performance work for large-model GPU deployments with an accent on benchmarking, profiling, distributed serving, and runtime optimization. Focus on diagnosing serving-stack bottlenecks, improving single-node and multi-node inference efficiency, and building production-ready runtimes and configurations.

Senior ML Systems Engineer Inference

Company

Runpod

Conditions

4 days agoSalary: 150K - 220K

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will lead end-to-end LLM serving performance work. You will define rigorous performance measurements, diagnose bottlenecks across the serving stack, improve efficiency for large models on GPU deployments, and turn findings into reliable production runtimes and configurations. You will collaborate on inference offerings, evaluate emerging ecosystem tools, and implement runtime fixes.

Requirements

  • 5+ years of professional system engineering experience
  • Production or serious benchmark-scale experience with vLLM, SGLang, or a comparable serving engine
  • Strong Python software engineering skills
  • Understanding of LLM inference performance, batching, memory, parallelism, latency, and throughput
  • Experience with quantization, speculative decoding, or distributed serving
  • Benchmarking, performance analysis, and GPU profiling skills
  • Clear written communication of results and decisions

Responsibilities

  • Define rigorous, repeatable inference-performance measurements
  • Profile and diagnose serving-stack performance problems
  • Improve serving efficiency for large models on single-node and multi-node GPU deployments
  • Create production-ready runtimes, configurations, and defaults
  • Shape inference offerings with product and infrastructure stakeholders
  • Evaluate, adopt, build, and contribute to inference ecosystem projects
  • Trace serving-engine bottlenecks and implement fixes

Benefits

  • Equity through stock options
  • Medical, dental, and vision plans
  • Flexible PTO
  • Remote-first work
  • $1,200 home office and equipment stipend

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →

Текст вакансии взят без изменений

Источник -