5 часов назад
Performance Engineer (AI)
200 000 - 400 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Performance Engineer (AI) (CUDA/GPU inference): Writing kernels and low-level optimizations that make vLLM faster across NVIDIA GPUs and emerging accelerator hardware with an accent on GPU architecture, profiling, and high-performance C++ and Python. Focus on optimizing ML-specific kernels, quantization, multi-platform accelerator support, and compiler technologies.
Location: San Francisco, California; remote work may be considered within the US for exceptional candidates
Salary: $200,000–$400,000 annual salary plus equity
Company
develops and advances vLLM as an AI inference engine, focusing on making model inference faster and more cost-efficient across modern hardware.
What you will do
- Write CUDA and other accelerator kernels for high-performance inference.
- Develop low-level optimizations to improve vLLM speed across NVIDIA GPUs and emerging accelerator platforms.
- Work directly with hardware vendors to integrate and optimize new chips.
- Profile workloads and benchmark performance using tools such as Nsight and rocprof.
- Optimize ML-specific kernels, including FlashAttention and fused kernels.
- Contribute to inference engine, GPU systems, and compiler optimization projects.
Requirements
- Bachelor's degree or equivalent experience in computer science, engineering, or a similar field.
- Deep experience writing CUDA kernels or equivalent kernels with CuTeDSL, Triton, TileLang, or Pallas.
- Strong understanding of GPU architecture, including memory hierarchy, warp scheduling, tiling, and tensor cores.
- Proficiency in C++ and Python, with demonstrated ability to write high-performance code.
- Experience with profiling tools and performance optimization methodologies.
- Ability to work onsite in San Francisco or remotely from the US if selected as an exceptional candidate.
Nice to have
- Knowledge of quantization techniques such as INT8, FP8, and mixed precision.
- Familiarity with NVIDIA, AMD, TPU, and Intel accelerator platforms.
- Experience with LLVM, MLIR, or XLA compiler technologies.
- Contributions to vLLM, other inference engines, GPU systems, or compiler optimization projects.
- Technical writing experience focused on GPU optimization.
Culture & Benefits
- Work at the intersection of AI models and hardware.
- Collaborate directly with hardware vendor teams.
- Health, dental, and vision benefits.
- 401(k) company match.
- Equity is included in the compensation package.
- Visa sponsorship is available on a case-by-case basis.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
5 дней назад
Sr. Inference Optimization Engineer (AI)
195 200 - 361 200$
8 часов назад
AI Algorithm Engineer (Hardware Co-Design)
200 000 - 400 000$
10 часов назад
AI Engineer (Simulation)
120 000 - 150 000$
7 часов назад
AI Platform Engineer
180 000 - 220 000$
10 часов назад
LLM Inference Deployment Engineer (AI)
180 000 - 240 000$
9 часов назад
AI Product Engineer (Semiconductor)
200 000 - 400 000$