3 часа назад
Technical Lead, On-Device AI Inference
300 000 - 500 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Technical Lead, On-Device AI Inference (AI hardware and multimodal models): Building the low-level inference stack that runs transformer models on Hark's custom silicon with an accent on accelerator selection, latency, memory, and power efficiency. Focus on designing execution layers, custom kernels, runtime systems, and compiler paths, while leading the team responsible for performance-critical on-device AI software.
Location: San Jose, United States
Salary: $300,000–$500,000 annual base salary
Company
is an artificial intelligence company developing personalized, proactive, multimodal intelligence and next-generation AI hardware.
What you will do
- Evaluate GPUs, NPUs, DSPs, and specialized accelerators for on-device model deployment.
- Co-design foundation model and audio ML architectures around latency, memory, and power constraints.
- Build low-level execution layers, custom kernels, runtime systems, and compiler paths for transformer workloads.
- Partner with silicon vendors and internal hardware teams to bring up new accelerators.
- Hire and lead engineers building performance-critical inference software.
Requirements
- 8–12+ years of experience in high-performance computing.
- Production experience deploying workloads on GPUs, NPUs, or specialized accelerators.
- Deep knowledge of attention, KV-cache behavior, quantization, and memory bandwidth limitations.
- Experience designing or optimizing inference engines, distributed runtimes, or ML compilers, including writing performance-critical kernels.
- Experience leading teams and setting technical direction for performance-critical software.
- Experience taking research models from checkpoints to constrained hardware used in a product.
Nice to have
- Experience with Hexagon DSP, Ambiq-class MCUs, or comparable embedded AI silicon.
- Experience with speech, audio, or streaming multimodal inference.
- Contributions to open-source inference or compiler toolchains such as TensorRT, ONNX Runtime, TVM, or MLIR.
Culture & Benefits
- Work on multimodal AI systems combining speech, text, vision, and persistent memory.
- Develop AI models and hardware together as a unified interface between humans and machines.
- Additional compensation components and benefits may be included depending on the role.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →