Назад
Company hidden
4 дня назад

MTS - Kernel Engineer (AI)

260 000 - 320 000$
Формат работы
remote (только USA)
Тип работы
fulltime
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
MTS - Kernel Engineer (AI): Developing and optimizing compute kernels for a custom AI accelerator with an accent on tensor operations, data movement, memory hierarchies, and profiling. Focus on defining kernel execution models, enabling simulation and pre-silicon execution, integrating with compiler infrastructure, and solving hardware-software performance bottlenecks.

Location: Mountain View, California, United States

Salary: $260,000–$320,000 per year

Company

hirify.global is hiring for an early-stage AI hardware company developing a custom accelerator platform.

What you will do

  • Develop and optimize compute kernels for a custom AI accelerator.
  • Implement tensor operations, data-movement patterns, and memory-hierarchy optimizations.
  • Build profiling infrastructure to measure throughput, latency, and performance against architectural targets.
  • Define execution, synchronization, register-passing, and memory-management strategies across control cores and specialized compute engines.
  • Enable kernel execution in simulation and pre-silicon environments and collaborate with compiler, architecture, and systems teams.
  • Investigate bottlenecks, recommend hardware or software improvements, and create technical documentation and development guides.

Requirements

  • Strong C and C++ programming skills in production-quality systems or performance-critical code.
  • Deep experience with CUDA or a comparable accelerator-programming model.
  • Strong knowledge of warp, wavefront, or thread-group execution, memory coalescing, shared or local memory, registers, caches, synchronization, and data locality.
  • Ability to reason about pipelines, execution units, memory hierarchies, bandwidth, latency, and data-movement costs.
  • Experience profiling and optimizing performance, including measuring throughput and latency and identifying bottlenecks.
  • Practical understanding of GEMM, convolution, attention, reductions, scatter/gather, and elementwise operations, plus Python for tooling and automation.

Nice to have

  • Experience with Triton, CUTLASS, MLIR, LLVM, kernel DSLs, or compiler infrastructure.
  • Knowledge of RISC-V, x86, ARM64, architectural simulators, or instruction-set simulators.
  • Experience with custom ASICs, GPUs, NPUs, FPGA development, RTL, Verilog, or SystemVerilog.
  • Background in high-performance or scientific computing and hardware-software co-design.

Culture & Benefits

  • Cross-functional collaboration with architecture, compiler, simulation, hardware, and systems teams.
  • Direct influence on the kernel programming model and the efficiency of custom accelerator hardware.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →