Sr. Inference Optimization Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
TL;DR
Sr. Inference Optimization Engineer (AI): Optimizing inference engines like llama.cpp and vLLM for constrained local and edge environments with an accent on latency, throughput, and memory efficiency. Focus on tuning KV cache, continuous batching, quantization strategies, and reducing CPU overhead to make hybrid agent products viable.
Location: Hybrid (US: Santa Clara, Phoenix, Folsom, or Hillsboro)
Salary: $195,200 - $361,200 USD
Company
A global leader in semiconductor manufacturing and AI hardware, focusing on making AI safer, more private, and sustainable through local and edge ligence.
What you will do
- Profile and optimize local inference engines (llama.cpp-vulkan and vLLM) for edge hardware latency and throughput.
- Tune KV cache, continuous batching, and scheduling for interactive agent workloads.
- Drive and validate quantization strategies using GGUF, AWQ, and GPTQ.
- Reduce CPU overhead and improve engine startup, model load, and lifecycle management.
- Benchmark performance across various hardware tiers and publish comparisons.
- Contribute fixes and patches to open-source inference engines.
Requirements
- BS/MS in CS, EE, Math, or a related STEM field.
- 8+ years of software development experience.
- Proficiency in C++ and/or Python with the ability to read systems-level code.
- Direct experience with LLM inference (attention, KV cache, decoding).
- Proven experience profiling and optimizing real-world performance problems on CPU or GPU.
- Expertise in Linux, build systems, and low-level debugging.
- Must be based in the United States.
Nice to have
- Hands-on experience with llama.cpp, vLLM, or ggml.
- Experience with GPU/accelerator programming (Vulkan, CUDA, SYCL, Metal) or SIMD/CPU kernels.
- Familiarity with quantization formats and their quality trade-offs.
- Previous open-source contributions to inference engines.
Culture & Benefits
- Total rewards package including competitive pay and stock bonuses.
- Comprehensive health, retirement, and vacation programs.
- Hybrid work model allowing a split between on-site and off-site work.
- Opportunity to work on the forefront of local/edge AI and hardware optimization.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →