Назад
Company hidden
14 часов назад

Member of Technical Staff, Hardware, Kernel Engineer (Custom Silicon)

200 000 - 420 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Member of Technical Staff, Hardware, Kernel Engineer (Custom Silicon) (C++/Custom Silicon): Building kernel generators and optimized low-level assembly for a high-performance custom silicon compute engine with an accent on deep learning primitives, memory hierarchy management, and hardware/software co-design. Focus on tuning instruction scheduling, register allocation, software pipelining, and data movement while validating performance against simulators and silicon.

Location: Austin, Texas or Palo Alto, California

Annual salary: $200,000–$420,000 USD

Company

hirify.global is developing personal AI through personal hardware for local inference, custom training infrastructure, next-generation user interfaces, and deep learning research.

What you will do

  • Design C++ code-generation frameworks and metaprogramming toolchains that emit optimized assembly for custom hardware.
  • Implement and optimize deep learning primitives including GEMM/MatMul, attention, convolutions, and element-wise layers.
  • Tune instruction scheduling, register allocation, and software pipelining to maximize hardware utilization and reduce execution latency.
  • Develop tiling, double-buffering, and data-movement strategies for efficient SRAM use and reduced memory bandwidth bottlenecks.
  • Collaborate with compiler, RTL, architecture, and deep learning teams on hardware/software co-design.
  • Benchmark and validate generated assembly using hardware simulators, silicon, and performance counters.

Requirements

  • Bachelor’s degree in Computer Engineering, Computer Science, Electrical Engineering, or a related field, plus 5+ years of practical industry experience in low-level performance programming.
  • Experience with hardware programming models such as CUDA, Triton, CUTLASS, or custom accelerator assembly, with a record of shipping optimized kernels.
  • Advanced knowledge of computer architecture, including vector units, execution pipelines, register files, caches, SRAM, HBM, and DRAM.
  • Proficiency in modern C++ for scalable metaprogramming and code-generation frameworks.
  • Strong mathematical foundation in linear algebra and deep learning primitives.
  • Visa sponsorship is available; the role is based in the United States.

Nice to have

  • Experience optimizing Tensor Cores, matrix multiply-accumulate units, or custom vector extensions.
  • Advanced C++ template metaprogramming or automated generation of parameterized kernel variants.
  • Experience with hardware profiling, execution tracing, performance counters, cache misses, pipeline stalls, and ALU bubbles.

Culture & Benefits

  • Work alongside scientists, engineers, and builders from leading technology companies and AI labs.
  • Health, dental, and vision benefits.
  • Unlimited paid time off.
  • Relocation support as needed.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →