7 часов назад
Inference Runtime Engineer (AI)
200 000 - 400 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Inference Runtime Engineer (AI): Optimizing vLLM inference execution for LLM and diffusion models across diverse hardware and architectures with an accent on transformer serving, model execution, and inference performance. Focus on implementing research-based inference techniques, managing KV-cache and hybrid serving, and solving complex performance challenges in ML codebases.
Location: On-site in San Francisco, California; remote work may be considered in the US for exceptional candidates.
Salary: $200,000–$400,000 annual salary plus equity
Company
develops and advances vLLM as an AI inference engine, focusing on making model inference faster and more cost-efficient across models and hardware.
What you will do
- Optimize how LLM and diffusion models execute across diverse hardware and model architectures.
- Develop and improve inference runtime capabilities in vLLM.
- Implement model architectures and inference techniques from research papers.
- Contribute performant, maintainable code to complex machine-learning codebases.
- Debug inference systems and support evolving architectures such as mixture-of-experts, multimodal, and agentic models.
Requirements
- Bachelor's degree or equivalent experience in computer science, engineering, or a related field.
- Deep understanding of transformer architectures and their variants.
- Strong Python programming skills and experience with PyTorch internals.
- Experience with LLM inference systems such as vLLM, TensorRT-LLM, SGLang, or TGI.
- Ability to understand and implement model architectures and inference techniques from research papers.
- Ability to write performant, maintainable code and debug complex ML codebases.
Nice to have
- Knowledge of KV-cache memory management, prefix caching, and hybrid model serving.
- Familiarity with reinforcement-learning frameworks and algorithms for LLMs.
- Experience with multimodal inference across audio, image, video, and text.
- Contributions to open-source machine-learning or systems-infrastructure projects.
- Experience contributing core features or integrations to vLLM and related inference projects.
Culture & Benefits
- Work at the intersection of AI models and hardware on the vLLM inference engine.
- Health, dental, and vision benefits.
- 401(k) with company match.
- Equity included in the compensation package.
- Visa sponsorship is available on a case-by-case basis.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
11 часов назад
LLM Inference Deployment Engineer (AI)
180 000 - 240 000$
11 часов назад
AI Engineer (LLM)
200 000 - 400 000$
5 дней назад
Sr. Inference Optimization Engineer (AI)
195 200 - 361 200$
Baseten
24 часа назад
Forward Deployed Engineers (AI)
200 000 - 400 000$
6 часов назад
AI Engineer (Fintech)
180 000 - 260 000$
11 часов назад
AI Research Engineer (Agentic LLMs)
200 000 - 400 000$