Назад
Company hidden
7 часов назад

Inference Runtime Engineer (AI)

200 000 - 400 000$
Формат работы
remote (только USA)/onsite
Тип работы
fulltime
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Inference Runtime Engineer (AI): Optimizing vLLM inference execution for LLM and diffusion models across diverse hardware and architectures with an accent on transformer serving, model execution, and inference performance. Focus on implementing research-based inference techniques, managing KV-cache and hybrid serving, and solving complex performance challenges in ML codebases.

Location: On-site in San Francisco, California; remote work may be considered in the US for exceptional candidates.

Salary: $200,000–$400,000 annual salary plus equity

Company

hirify.global develops and advances vLLM as an AI inference engine, focusing on making model inference faster and more cost-efficient across models and hardware.

What you will do

  • Optimize how LLM and diffusion models execute across diverse hardware and model architectures.
  • Develop and improve inference runtime capabilities in vLLM.
  • Implement model architectures and inference techniques from research papers.
  • Contribute performant, maintainable code to complex machine-learning codebases.
  • Debug inference systems and support evolving architectures such as mixture-of-experts, multimodal, and agentic models.

Requirements

  • Bachelor's degree or equivalent experience in computer science, engineering, or a related field.
  • Deep understanding of transformer architectures and their variants.
  • Strong Python programming skills and experience with PyTorch internals.
  • Experience with LLM inference systems such as vLLM, TensorRT-LLM, SGLang, or TGI.
  • Ability to understand and implement model architectures and inference techniques from research papers.
  • Ability to write performant, maintainable code and debug complex ML codebases.

Nice to have

  • Knowledge of KV-cache memory management, prefix caching, and hybrid model serving.
  • Familiarity with reinforcement-learning frameworks and algorithms for LLMs.
  • Experience with multimodal inference across audio, image, video, and text.
  • Contributions to open-source machine-learning or systems-infrastructure projects.
  • Experience contributing core features or integrations to vLLM and related inference projects.

Culture & Benefits

  • Work at the intersection of AI models and hardware on the vLLM inference engine.
  • Health, dental, and vision benefits.
  • 401(k) with company match.
  • Equity included in the compensation package.
  • Visa sponsorship is available on a case-by-case basis.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →