Назад
Company hidden
4 часа назад

On-Device AI Inference Engineer (Embedded AI)

200 000 - 450 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
On-Device AI Inference Engineer (Embedded AI): Making multimodal transformer models run efficiently on constrained hardware with an accent on low-level kernels, runtime paths, profiling, and quantization. Focus on optimizing memory, latency, and power across DSPs, NPUs, and emerging accelerators while integrating deployment constraints into model architecture decisions.

Location: San Jose, United States

Salary: $200,000–$450,000 base salary annually

Company

hirify.global is an artificial intelligence company developing personalized multimodal intelligence and next-generation hardware interfaces that combine speech, text, vision, and persistent memory.

What you will do

  • Write and optimize low-level kernels and runtime paths for transformer workloads on target silicon.
  • Design model residency, scheduling, and swap behavior across concurrent workloads sharing limited memory and power.
  • Profile models on real hardware, identify bottlenecks, and improve delivered performance.
  • Convert models from full precision to INT8 and INT4 while meeting product size, latency, and power budgets.
  • Optimize transformer workloads for new DSPs, NPUs, and other accelerators in collaboration with the hardware team.
  • Provide deployment constraints to model teams so architecture decisions reflect hardware capabilities.

Requirements

  • 4–8+ years of experience writing performance-critical software and optimizing workloads on GPUs, NPUs, DSPs, or similar accelerators.
  • Strong C and C++ skills, including SIMD, custom kernels, memory layout, and profiling tools.
  • Understanding of attention, KV-cache behavior, transformer inference, and memory bandwidth.
  • Experience optimizing against compute, memory, and power budgets.
  • Experience deploying an optimized model in a product running on constrained hardware.

Nice to have

  • Experience with Hexagon DSP, Ambiq-class MCUs, or comparable embedded AI silicon.
  • Familiarity with ONNX Runtime, TVM, MLIR, TensorRT, or similar inference and compiler toolchains.
  • Background in speech or audio inference.

Culture & Benefits

  • Hands-on systems work close to the hardware.
  • Small-team environment with direct impact on product response time.
  • Work alongside hardware and model teams to develop unified AI systems and devices.
  • Full-time position with a US base salary and potential additional compensation and benefits.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →