Назад
3 месяца назад

Embedded AI Engineer (Edge AI)

219 300 - 274 100$
Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior/lead
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Embedded AI Engineer (Edge AI): Developing and optimizing custom kernels, operators, and runtime components that bring Deepgram speech models to non-NVIDIA accelerators, embedded SoCs, DSPs, and NPUs with an accent on quantization, architecture-specific compilation, and low-level performance tuning. Focus on meeting latency, memory, power, and thermal constraints, integrating vendor toolchains, and validating production inference across constrained hardware.

Location: USA — Remote

Base salary: $219.3K–$274.1K annually, excluding bonus, equity, and benefits.

Company

Deepgram provides real-time speech-to-text, text-to-speech, and voice AI APIs and deployable foundation models for large-scale production use.

What you will do

  • Write and optimize custom kernels and operators in C, C++, Rust, and platform-specific assembly or intrinsics.
  • Optimize speech models for non-NVIDIA accelerators, embedded SoCs, DSPs, NPUs, and purpose-built devices.
  • Apply quantization, operator fusion, memory-layout optimization, and architecture-specific compilation.
  • Integrate vendor NPU/DSP toolchains and edge inference runtimes, extending them with custom operators.
  • Build runtime components for embedded Linux, bare-metal, and RTOS environments.
  • Benchmark and validate latency, accuracy, power, memory footprint, utilization, and regressions across target platforms.

Requirements

  • Production experience with resource-constrained embedded, mobile, or edge AI systems.
  • Strong proficiency in C, C++, and/or Rust for performance-critical environments.
  • Hands-on experience with on-device model optimization, including quantization, pruning, knowledge distillation, or architecture-specific compilation.
  • Familiarity with edge inference runtimes such as ONNX Runtime, TensorRT, TFLite, or ExecuTorch, and/or vendor NPU/DSP toolchains.
  • Understanding of CPU, GPU, NPU, and DSP architectures, memory hierarchies, fixed-point arithmetic, and power management.
  • Experience with bare-metal, RTOS, embedded Linux, microcontrollers, or edge SoC development.

Nice to have

  • Real-time audio processing, DSP pipelines, audio codec optimization, wake-word detection, or streaming inference.
  • Custom quantization, mixed-precision inference, neural architecture search, or model compilation toolchains.
  • Hardware evaluation and benchmarking across accelerators, SoCs, or GPUs.
  • Experience shipping AI features in consumer products or implementing secure on-device deployment.

Culture & Benefits

  • AI-first operating model with active use and experimentation with advanced AI tools.
  • Fast-changing environment focused on experimentation, adaptation, and continuous learning.
  • Compensation includes annual bonus and equity in addition to base salary.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →