Назад
Company hidden
обновлено 4 дня назад

Inference Performance Engineer (AI)

Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Inference Performance Engineer (AI): Optimizing the inference stack for efficient model serving with an accent on caching, batching, quantization, decoding, and kernel-level performance. Focus on improving throughput, cost, and tail latency, tuning routing across infrastructure and external providers, and building profiling systems for production workloads.

Location: San Francisco; hybrid work with in-person collaboration in the Bay Area

Company

Builds AI systems designed to adapt in real time through efficient, flexible, personalized, and accessible intelligence.

What you will do

  • Own the cost and performance of the model inference stack as workloads, traffic, and hardware change.
  • Improve throughput, cost, and tail latency through KV-cache management, continuous batching, speculative decoding, and quantization.
  • Optimize long-context prefill and decode workloads using real production traffic.
  • Tune routing between internal infrastructure and external providers based on cost, capacity, and performance.
  • Optimize serving engines such as vLLM, SGLang, and TensorRT-LLM, including below-framework work when needed.
  • Build profiling and measurement systems to identify time, memory, and compute usage.

Requirements

  • 5+ years of experience in ML systems, inference infrastructure, or performance engineering, with measurable cost or latency improvements.
  • Deep understanding of model serving, including prefill and decode, memory bandwidth, batching, and concurrency.
  • Production experience with vLLM, SGLang, or TensorRT-LLM.
  • Strong Python skills and proficiency in C++, Rust, or another systems language.
  • Experience with GPU performance, including CUDA, NCCL, mixed precision, memory layout, kernels, or quantization.

Culture & Benefits

  • In-person collaboration in the Bay Area within a distributed, global-first team.
  • Team offsites and an annual travel stipend for exploring a new country.
  • Weekly meal allowance for takeout or grocery delivery.
  • Comprehensive medical benefits and generous paid time off.
  • Collaborative environment encouraging adaptable teammates and bold ideas.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →