Назад
Company hidden
8 дней назад

GPU Performance Engineer (Machine Learning)

Формат работы
onsite
Тип работы
fulltime
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
GPU Performance Engineer (CUDA/Machine Learning): Building highly optimized CUDA kernels and production-grade GPU implementations for low-latency inference across neural networks, tree-based models, and other structured workloads with an accent on kernel optimization, memory layouts, and hardware-aware execution. Focus on profiling and benchmarking GPU performance, improving inference latency and throughput, and translating quantitative models into efficient compute pipelines.

Location: Onsite in Bala Cynwyd, Philadelphia Area, Pennsylvania, United States

Company

hirify.global is a global quantitative trading firm using machine learning, scientific research, and advanced technology to develop systematic trading strategies.

What you will do

  • Design, implement, and optimize custom CUDA kernels for latency-critical inference workloads.
  • Develop fine-grained GPU implementations tailored to neural networks, tree-based models, and other structured model architectures.
  • Analyze quantitative research models and computational bottlenecks to identify parallelization and hardware-efficiency opportunities.
  • Translate mathematical models into production-grade, high-performance compute pipelines in collaboration with quantitative researchers.
  • Optimize inference through kernel tuning, memory-layout design, execution strategies, I/O optimization, and precision tradeoffs.
  • Profile and benchmark GPU performance, improve production latency and throughput, and contribute to GPU architecture decisions.

Requirements

  • Strong proficiency in writing and optimizing CUDA kernels.
  • Solid programming experience in C/C++.
  • Deep understanding of GPU architecture, including memory hierarchy, SIMT execution, occupancy, and latency/throughput tradeoffs.
  • Ability to reason about numerical stability, precision, performance tradeoffs, and hardware-efficient model design.
  • Strong problem-solving skills and comfort working with low-level systems.

Nice to have

  • PhD in mathematics, physics, computer science, engineering, or a related quantitative field.
  • Background in linear algebra, probability, numerical methods, or scientific computing.
  • Experience with quantitative research teams, financial models, or real-world inference optimization beyond baseline frameworks and libraries.
  • Familiarity with PTX-level behavior, tensor cores, architecture-specific tuning, ONNX Runtime, TensorRT, Triton, TVM, or similar systems.
  • Experience with neural networks, LightGBM, Mamba architectures, kernel fusion, custom operators, model compilation, or graph-level optimization.

Culture & Benefits

  • Intellectually driven and highly collaborative quantitative trading environment.
  • Close collaboration among researchers, engineers, and traders.
  • Opportunity to solve complex problems involving global markets, machine learning, and advanced quantitative research.
  • Immediate start availability.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →