Назад
Company hidden
5 часов назад

CUDA Kernel Engineer

Формат работы
onsite
Тип работы
fulltime
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
CUDA Kernel Engineer (CUDA/GPU): Developing and optimizing state-of-the-art CUDA kernels for AI models used in semiconductor design and verification, with an accent on large-scale training, inference, and reinforcement learning workloads. Focus on profiling GPU performance, integrating custom kernels with AI frameworks, and building GPU-accelerated primitives for graph reasoning, symbolic computation, and hardware simulation.

Location: Palo Alto Office, United States

Company

hirify.global develops world models and AI agents for understanding and building hardware, electronics systems, and semiconductors.

What you will do

  • Develop, integrate, and optimize CUDA kernels for AI models used in semiconductor design and verification.
  • Improve large-scale model training, inference, and reinforcement learning systems running across thousands of GPUs.
  • Build performance tools, benchmarks, and integration layers to maximize GPU utilization.
  • Develop GPU-accelerated primitives for graph reasoning, symbolic computation, and hardware simulation.
  • Collaborate with AI researchers and semiconductor experts to translate domain-specific workloads into high-performance GPU code.
  • Release kernels and tooling as contributions to open-source AI and HPC ecosystems.

Requirements

  • Experience writing and optimizing CUDA kernels for large-scale AI workloads.
  • Experience profiling and optimizing GPU performance for custom compute- or memory-bound workloads.
  • Experience integrating custom kernels with training and inference frameworks such as PyTorch, Megatron, vLLM, or TorchTitan.
  • Experience with current NVIDIA hardware and software stacks, including Hopper, Blackwell, NVLink, NCCL, or Triton.
  • Ability to collaborate with AI researchers and semiconductor specialists on high-performance GPU implementations.

Culture & Benefits

  • Work alongside researchers, engineers, and semiconductor experts from leading academic and technology organizations.
  • Contribute to open-source AI and high-performance computing ecosystems.
  • Work on AI systems designed to reason about circuit layouts, generate and validate RTL, and optimize chip architectures.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →