Назад
Company hidden
4 дня назад

GPU Performance Engineer (CUDA/Triton)

Формат работы
hybrid
Тип работы
fulltime
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
GPU Performance Engineer (CUDA/Triton) (Multimodal AI): Profiling and optimizing GPU, CPU, and accelerator code for multimodal model training and large-scale deployment with an accent on fused kernels, tensor cores, and transformer internals. Focus on building high-performance PyTorch, Triton, and CUDA operations, optimizing distributed multi-node systems, and preventing performance regressions through monitoring and automation.

Location: Redwood City, California, United States; hybrid

Company

hirify.global builds unified general intelligence systems that generate, understand, and operate in the physical world, with a focus on multimodal AI and vision.

What you will do

  • Profile and optimize GPU, CPU, and accelerator code to maximize utilization and minimize latency.
  • Develop high-performance PyTorch, Triton, and CUDA kernels and custom operations.
  • Build fused kernels using tensor cores and modern hardware features across platforms.
  • Optimize transformer model architectures and implementations for distributed multi-node production deployment.
  • Create performance monitoring, analysis, and automation tools.
  • Research and implement optimization techniques for transformer models.

Requirements

  • Expert-level Triton and CUDA programming with strong GPU optimization skills.
  • Strong PyTorch experience, including kernel development and custom operations.
  • Proficiency with profiling tools such as NVIDIA Nsight, torch profiler, and custom tooling.
  • Deep understanding of transformer architectures and attention mechanisms.
  • Ability to work in a hybrid setup in Redwood City, California.

Nice to have

  • Experience with compilers and exporters such as torch.compile, TensorRT, ONNX, or XLA.
  • Experience optimizing inference workloads for latency and throughput.
  • Knowledge of Triton compiler and kernel fusion techniques.
  • Knowledge of warp-level intrinsics and advanced CUDA optimization.

Culture & Benefits

  • Work on multimodal models intended to support general intelligence.
  • Equal opportunity employment environment.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →