4 часа назад
On-Device AI Inference Engineer (Embedded AI)
200 000 - 450 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
On-Device AI Inference Engineer (Embedded AI): Making multimodal transformer models run efficiently on constrained hardware with an accent on low-level kernels, runtime paths, profiling, and quantization. Focus on optimizing memory, latency, and power across DSPs, NPUs, and emerging accelerators while integrating deployment constraints into model architecture decisions.
Location: San Jose, United States
Salary: $200,000–$450,000 base salary annually
Company
is an artificial intelligence company developing personalized multimodal intelligence and next-generation hardware interfaces that combine speech, text, vision, and persistent memory.
What you will do
- Write and optimize low-level kernels and runtime paths for transformer workloads on target silicon.
- Design model residency, scheduling, and swap behavior across concurrent workloads sharing limited memory and power.
- Profile models on real hardware, identify bottlenecks, and improve delivered performance.
- Convert models from full precision to INT8 and INT4 while meeting product size, latency, and power budgets.
- Optimize transformer workloads for new DSPs, NPUs, and other accelerators in collaboration with the hardware team.
- Provide deployment constraints to model teams so architecture decisions reflect hardware capabilities.
Requirements
- 4–8+ years of experience writing performance-critical software and optimizing workloads on GPUs, NPUs, DSPs, or similar accelerators.
- Strong C and C++ skills, including SIMD, custom kernels, memory layout, and profiling tools.
- Understanding of attention, KV-cache behavior, transformer inference, and memory bandwidth.
- Experience optimizing against compute, memory, and power budgets.
- Experience deploying an optimized model in a product running on constrained hardware.
Nice to have
- Experience with Hexagon DSP, Ambiq-class MCUs, or comparable embedded AI silicon.
- Familiarity with ONNX Runtime, TVM, MLIR, TensorRT, or similar inference and compiler toolchains.
- Background in speech or audio inference.
Culture & Benefits
- Hands-on systems work close to the hardware.
- Small-team environment with direct impact on product response time.
- Work alongside hardware and model teams to develop unified AI systems and devices.
- Full-time position with a US base salary and potential additional compensation and benefits.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →