Назад
Company hidden
2 дня назад

ML Systems Engineer (Inference Infrastructure)

Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
ML Systems Engineer (Inference Infrastructure): Own and optimize the cost and performance of the inference stack, focusing on caching, batching, quantization, decoding, and kernel-level optimization. Focus on improving throughput, latency, and reliability while working with serving engines like vLLM, SGLang, and TensorRT-LLM.

Location

Location: San Francisco, hybrid work format

Company

hirify.global is a startup focused on building efficient, adaptable AI systems that evolve in real-time to expand access and innovation.

What you will do

  • Own cost and performance of the inference stack, optimizing caching, batching, quantization, decoding, and kernel-level operations.
  • Improve throughput, cost, and tail latency through KV-cache management, continuous batching, speculative decoding, and quantization.
  • Optimize long-context prefill and decode workloads based on real production traffic.
  • Tune routing between infrastructure and external providers based on cost, capacity, and performance.
  • Work within serving engines such as vLLM, SGLang, and TensorRT-LLM, including low-level framework optimizations.
  • Build profiling and measurement systems to analyze time, memory, and compute usage.

Requirements

  • 5+ years experience in ML systems, inference infrastructure, or performance engineering with measurable improvements in cost or latency.
  • Deep understanding of model serving, including prefill, decode, memory bandwidth, batching, and concurrency.
  • Production experience with serving engines such as vLLM, SGLang, or TensorRT-LLM.
  • Strong Python skills and proficiency in C++, Rust, or another systems language.
  • Experience with GPU performance including CUDA, NCCL, mixed precision, memory layout, kernels, or quantization.

Culture & Benefits

  • Flexible work with in-person collaboration in the Bay Area and a distributed global-first team.
  • Annual travel stipend called hirify.global Passport to explore new countries.
  • Weekly lunch stipend for take-out or grocery delivery.
  • Comprehensive medical benefits and generous paid time off.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →