Назад
Company hidden
5 часов назад

Performance Engineer (AI)

200 000 - 400 000$
Формат работы
remote (только USA)/onsite
Тип работы
fulltime
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Performance Engineer (AI) (CUDA/GPU inference): Writing kernels and low-level optimizations that make vLLM faster across NVIDIA GPUs and emerging accelerator hardware with an accent on GPU architecture, profiling, and high-performance C++ and Python. Focus on optimizing ML-specific kernels, quantization, multi-platform accelerator support, and compiler technologies.

Location: San Francisco, California; remote work may be considered within the US for exceptional candidates

Salary: $200,000–$400,000 annual salary plus equity

Company

hirify.global develops and advances vLLM as an AI inference engine, focusing on making model inference faster and more cost-efficient across modern hardware.

What you will do

  • Write CUDA and other accelerator kernels for high-performance inference.
  • Develop low-level optimizations to improve vLLM speed across NVIDIA GPUs and emerging accelerator platforms.
  • Work directly with hardware vendors to integrate and optimize new chips.
  • Profile workloads and benchmark performance using tools such as Nsight and rocprof.
  • Optimize ML-specific kernels, including FlashAttention and fused kernels.
  • Contribute to inference engine, GPU systems, and compiler optimization projects.

Requirements

  • Bachelor's degree or equivalent experience in computer science, engineering, or a similar field.
  • Deep experience writing CUDA kernels or equivalent kernels with CuTeDSL, Triton, TileLang, or Pallas.
  • Strong understanding of GPU architecture, including memory hierarchy, warp scheduling, tiling, and tensor cores.
  • Proficiency in C++ and Python, with demonstrated ability to write high-performance code.
  • Experience with profiling tools and performance optimization methodologies.
  • Ability to work onsite in San Francisco or remotely from the US if selected as an exceptional candidate.

Nice to have

  • Knowledge of quantization techniques such as INT8, FP8, and mixed precision.
  • Familiarity with NVIDIA, AMD, TPU, and Intel accelerator platforms.
  • Experience with LLVM, MLIR, or XLA compiler technologies.
  • Contributions to vLLM, other inference engines, GPU systems, or compiler optimization projects.
  • Technical writing experience focused on GPU optimization.

Culture & Benefits

  • Work at the intersection of AI models and hardware.
  • Collaborate directly with hardware vendor teams.
  • Health, dental, and vision benefits.
  • 401(k) company match.
  • Equity is included in the compensation package.
  • Visa sponsorship is available on a case-by-case basis.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →