Назад
1 час назад

Inference Performance Engineer (LLM Inference)

180 000 - 360 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US/Canada
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Inference Performance Engineer (LLM Inference): Building and optimizing production inference systems for demanding AI workloads with an accent on runtime internals, scheduling, routing, and GPU efficiency. Focus on profiling end-to-end latency, implementing quantization and speculative decoding, tuning new model architectures and hardware, and developing benchmarks for cost-effective model serving.

Location: Hybrid in San Francisco, Montreal, New York, Seattle, or Toronto

Salary: $180K–$360K annually, plus equity

Company

Baseten provides inference infrastructure and developer tooling that help AI companies deploy cutting-edge models in production.

What you will do

  • Implement and productionize inference techniques including quantization, speculative decoding, KV-cache reuse, chunked prefill, LoRA, and guided generation.
  • Profile and optimize inference across runtime internals, GPU kernels, memory layout, scheduling, batching, routing, and prefill/decode disaggregation.
  • Improve tokens per GPU-hour, utilization, latency, throughput, and serving costs.
  • Bring up and tune new model architectures on emerging hardware.
  • Build benchmarking frameworks across model architectures, batch sizes, sequence lengths, and hardware configurations.
  • Contribute to open-source inference engines and collaborate with model, infrastructure, and customer-facing teams.

Requirements

  • Bachelor's, Master's, or Ph.D. in Computer Science, Engineering, Mathematics, or a related field.
  • Experience with Python, C++, or another general-purpose programming language.
  • Familiarity with LLM optimization techniques such as quantization, speculative decoding, and continuous batching.
  • Strong familiarity with ML libraries, especially PyTorch, TensorRT, or TensorRT-LLM.
  • Demonstrated interest and experience in LLMs.
  • Deep understanding of GPU architecture.

Nice to have

  • Contributions to vLLM, SGLang, TensorRT-LLM, or another inference engine.
  • Experience with large-scale distributed serving, autoscaling, load balancing, or multi-region and multi-cloud capacity.
  • Experience optimizing GPU kernels with CUDA, Triton, CUTLASS, or similar technologies.
  • Production experience with FP8/FP4, AWQ, GPTQ, or speculative decoding.
  • Experience developing and deploying AI/ML inference solutions.

Culture & Benefits

  • Competitive compensation with meaningful equity.
  • Flexible PTO and a company-wide winter break.
  • Paid parental leave.
  • Fertility and family-building stipend through Carrot.
  • U.S. employees receive company-facilitated 401(k) and full medical, dental, and vision coverage for employees and dependents.
  • Exposure to a range of ML startups and opportunities for learning and networking.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →