Назад
Company hidden
2 месяца назад

AMD GPU Performance Engineer

200 000 - 400 000SGD
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Singapore
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
AMD GPU Performance Engineer (ROCm/HIP/vLLM): Building and optimizing AMD GPU backends, kernels, runtime paths, and benchmarking infrastructure for vLLM with an accent on inference performance across the AMD accelerator ecosystem. Focus on optimizing attention, GEMM, sampling, KV cache, and communication-heavy operations through profiling, hardware counters, correctness testing, and reproducible benchmarks.

Location: On-site in Singapore

Salary: S$200,000–S$400,000 annually, plus equity

Company

hirify.global develops vLLM as an AI inference engine, focusing on making model inference faster and more cost-efficient.

What you will do

  • Build and optimize AMD GPU backends, kernels, runtime paths, and benchmarking infrastructure for vLLM.
  • Use ROCm, HIP, Triton, CK, AITER, and related tools to improve inference performance on AMD GPUs.
  • Optimize performance-critical inference paths including attention, GEMM, sampling, KV cache, fused kernels, and communication-heavy operations.
  • Profile and benchmark workloads using measurements, hardware counters, correctness tests, and reproducible benchmarks.
  • Improve AMD GPU support in vLLM so it is usable, fast, benchmarked, and maintainable.

Requirements

  • Bachelor’s degree or equivalent experience in computer science, engineering, systems, machine learning, or a related field.
  • Hands-on experience optimizing AMD GPU workloads with ROCm, HIP, Triton, CK, AITER, or similar tools.
  • Deep understanding of AMD GPU execution, memory behavior, toolchains, kernel performance, and backend-specific constraints.
  • Experience optimizing ML kernels or inference paths such as attention, GEMM, sampling, KV cache, fused kernels, or communication-heavy runtime paths.
  • Strong performance profiling and benchmarking skills.
  • Ability to work on-site in Singapore.

Nice to have

  • Experience with vLLM, SGLang, TensorRT-LLM, ROCm-based serving, or other LLM inference systems.
  • Familiarity with batching, KV cache, decoding, serving trade-offs, and production inference systems.
  • Experience with Triton, MLIR, LLVM, CK, AITER, HIP, compiler projects, or kernel DSLs.
  • Knowledge of INT8, FP8, mixed precision, or AMD-specific numeric formats.
  • Open-source contributions or experience building AMD GPU benchmarking and regression detection infrastructure.

Culture & Benefits

  • Work at the intersection of inference systems, kernels, compilers, and hardware architecture.
  • Medical, dental, and vision coverage.
  • Visa sponsorship is available on a case-by-case basis.
  • Equity included in the compensation package.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →