Назад
Company hidden
5 дней назад

Senior GPU Performance Software Engineer (AI)

195 200 - 275 580$
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior GPU Performance Software Engineer (AI): Developing and optimizing oneDNN GPU kernels, math primitives, and code-generation infrastructure for deep learning frameworks with an accent on GEMM, convolution, attention, mixed-precision execution, and hardware utilization. Focus on profiling performance bottlenecks, designing scalable parallel algorithms, and co-designing GPU primitives with hardware and compiler teams.

Location: Hybrid work model in the United States, with on-site work at an assigned Intel site in Hillsboro, Oregon or Santa Clara, California.

Annual salary: $195,200–$275,580 USD

Company

hirify.global's Software and AI organization contributes to oneDNN, an open-source performance library that powers deep learning applications on Intel hardware.

What you will do

  • Develop high-performance GEMM, convolution, and attention kernels for AI workloads.
  • Design JIT and code-generation infrastructure for GPU kernel generation.
  • Implement fusion, memory-traffic, mixed-precision, and quantized execution optimizations.
  • Build performance models, profile GPU primitives and runtime paths, and eliminate bottlenecks.
  • Co-design GPU primitives and kernel architectures with hardware and compiler teams.
  • Improve validation, benchmarking, and CI infrastructure for performance-critical workloads.

Requirements

  • BSc, MSc, or PhD in Computer Science, Computer Engineering, Mathematics, Physics, or a related technical field.
  • 5+ years of professional software development experience with expert-level modern C++.
  • 2+ years of GPU programming and kernel optimization with SYCL/DPC++, OpenCL, CUDA, or HIP, or 5+ years of comparable low-level CPU performance optimization.
  • Strong knowledge of computer architecture, cache hierarchies, memory subsystems, and parallel programming, including multithreading and SIMD/vectorization.
  • Strong ownership, collaboration, and performance-focused problem-solving skills.

Nice to have

  • Experience developing high-performance math libraries, including GEMM, convolution, reduction, or FFT kernels.
  • GPU assembly-level tuning or compiler optimization experience.
  • Familiarity with OpenMP or oneTBB.
  • Basic understanding of deep learning primitives and upstream framework usage.

Culture & Benefits

  • Work on a high-impact open-source library used across millions of devices.
  • Collaborate with experts in GPU compilers, hardware architecture, and performance libraries.
  • Access and influence next-generation discrete GPU software stacks.
  • Competitive pay, stock programs, quarterly bonuses, healthcare, retirement benefits, and vacation.
  • Hybrid flexibility with time split between the assigned site and off-site work.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →