2 месяца назад
AMD GPU Performance Engineer
200 000 - 400 000SGD
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
AMD GPU Performance Engineer (ROCm/HIP/vLLM): Building and optimizing AMD GPU backends, kernels, runtime paths, and benchmarking infrastructure for vLLM with an accent on inference performance across the AMD accelerator ecosystem. Focus on optimizing attention, GEMM, sampling, KV cache, and communication-heavy operations through profiling, hardware counters, correctness testing, and reproducible benchmarks.
Location: On-site in Singapore
Salary: S$200,000–S$400,000 annually, plus equity
Company
develops vLLM as an AI inference engine, focusing on making model inference faster and more cost-efficient.
What you will do
- Build and optimize AMD GPU backends, kernels, runtime paths, and benchmarking infrastructure for vLLM.
- Use ROCm, HIP, Triton, CK, AITER, and related tools to improve inference performance on AMD GPUs.
- Optimize performance-critical inference paths including attention, GEMM, sampling, KV cache, fused kernels, and communication-heavy operations.
- Profile and benchmark workloads using measurements, hardware counters, correctness tests, and reproducible benchmarks.
- Improve AMD GPU support in vLLM so it is usable, fast, benchmarked, and maintainable.
Requirements
- Bachelor’s degree or equivalent experience in computer science, engineering, systems, machine learning, or a related field.
- Hands-on experience optimizing AMD GPU workloads with ROCm, HIP, Triton, CK, AITER, or similar tools.
- Deep understanding of AMD GPU execution, memory behavior, toolchains, kernel performance, and backend-specific constraints.
- Experience optimizing ML kernels or inference paths such as attention, GEMM, sampling, KV cache, fused kernels, or communication-heavy runtime paths.
- Strong performance profiling and benchmarking skills.
- Ability to work on-site in Singapore.
Nice to have
- Experience with vLLM, SGLang, TensorRT-LLM, ROCm-based serving, or other LLM inference systems.
- Familiarity with batching, KV cache, decoding, serving trade-offs, and production inference systems.
- Experience with Triton, MLIR, LLVM, CK, AITER, HIP, compiler projects, or kernel DSLs.
- Knowledge of INT8, FP8, mixed precision, or AMD-specific numeric formats.
- Open-source contributions or experience building AMD GPU benchmarking and regression detection infrastructure.
Culture & Benefits
- Work at the intersection of inference systems, kernels, compilers, and hardware architecture.
- Medical, dental, and vision coverage.
- Visa sponsorship is available on a case-by-case basis.
- Equity included in the compensation package.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
6 дней назад
Software Engineer, GPU Performance (AI)
6 дней назад
Research Engineer (Robotics)
18 часов назад
Senior Associate, Data Analytics and Machine Learning
6 дней назад
Senior Staff / Staff Machine Learning Engineer (AI/LLM)
6 дней назад
College Intern - Pen Health & Servicing Algorithm Development Engineer (Inkjet)
13 часов назад