Назад
Company hidden
2 часа назад

CUDA Engineer (AI)

Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
UK/US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
CUDA Engineer (AI): Writing and optimising low-level GPU code for transformer inference workloads with an accent on custom CUDA kernels, memory hierarchies, and performance profiling. Focus on kernel fusion, quantisation-aware and mixed-precision computation, autoregressive decoding, and maximising throughput across GPU architectures.

Location: Remote; London, England, United Kingdom and United States

Company

hirify.global is an energy startup developing integrated energy generation, hardware, grid infrastructure, real-time power trading, distributed energy systems, and high-performance compute infrastructure for AI workloads.

What you will do

  • Write and optimise custom CUDA kernels for transformer inference operations.
  • Profile kernels and eliminate occupancy, memory throughput, and warp divergence bottlenecks.
  • Apply kernel fusion, optimise memory access patterns, and manage GPU memory hierarchies.
  • Implement quantisation-aware and mixed-precision kernels, plus caching for autoregressive decoding.
  • Tune kernel launch configurations and benchmark performance against existing baselines.
  • Write regression tests, maintain internal CUDA libraries, and contribute to coding standards and documentation.

Requirements

  • 4+ years of production CUDA development experience with shipped performance-critical kernels.
  • Deep understanding of GPU microarchitecture, including warps, occupancy, register pressure, and memory hierarchy.
  • Strong CUDA C++ skills, including streams and asynchronous execution.
  • Hands-on experience profiling compute-bound and memory-bound bottlenecks.
  • Experience with kernel fusion, memory coalescing, warp-divergence avoidance, quantised kernels, and mixed-precision arithmetic.
  • Strong understanding of parallel algorithm design and numerical precision trade-offs.

Nice to have

  • Experience with transformer or attention kernels and autoregressive decoding.
  • Experience building high-performance GPU libraries from scratch.
  • HPC, latency-critical performance engineering, or multi-GPU and multi-node kernel optimisation experience.
  • Ability to read PTX/SASS to validate kernel efficiency.

Culture & Benefits

  • Remote work with global hiring; benefits vary by location.
  • Competitive salary with equity eligibility.
  • Biannual bonus scheme.
  • Fully expensed technology matched to work needs.
  • Private health insurance.
  • Breakfast and dinner allowance for office-based employees.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →