Назад
Company hidden
3 часа назад

Technical Lead, On-Device AI Inference

300 000 - 500 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
lead
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Technical Lead, On-Device AI Inference (AI hardware and multimodal models): Building the low-level inference stack that runs transformer models on Hark's custom silicon with an accent on accelerator selection, latency, memory, and power efficiency. Focus on designing execution layers, custom kernels, runtime systems, and compiler paths, while leading the team responsible for performance-critical on-device AI software.

Location: San Jose, United States

Salary: $300,000–$500,000 annual base salary

Company

hirify.global is an artificial intelligence company developing personalized, proactive, multimodal intelligence and next-generation AI hardware.

What you will do

  • Evaluate GPUs, NPUs, DSPs, and specialized accelerators for on-device model deployment.
  • Co-design foundation model and audio ML architectures around latency, memory, and power constraints.
  • Build low-level execution layers, custom kernels, runtime systems, and compiler paths for transformer workloads.
  • Partner with silicon vendors and internal hardware teams to bring up new accelerators.
  • Hire and lead engineers building performance-critical inference software.

Requirements

  • 8–12+ years of experience in high-performance computing.
  • Production experience deploying workloads on GPUs, NPUs, or specialized accelerators.
  • Deep knowledge of attention, KV-cache behavior, quantization, and memory bandwidth limitations.
  • Experience designing or optimizing inference engines, distributed runtimes, or ML compilers, including writing performance-critical kernels.
  • Experience leading teams and setting technical direction for performance-critical software.
  • Experience taking research models from checkpoints to constrained hardware used in a product.

Nice to have

  • Experience with Hexagon DSP, Ambiq-class MCUs, or comparable embedded AI silicon.
  • Experience with speech, audio, or streaming multimodal inference.
  • Contributions to open-source inference or compiler toolchains such as TensorRT, ONNX Runtime, TVM, or MLIR.

Culture & Benefits

  • Work on multimodal AI systems combining speech, text, vision, and persistent memory.
  • Develop AI models and hardware together as a unified interface between humans and machines.
  • Additional compensation components and benefits may be included depending on the role.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →