Назад
13 часов назад

Applied ML Engineer (Edge AI)

155 000 - 245 000$
Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Applied ML Engineer (Edge AI): Porting Deepgram speech models to non-NVIDIA accelerators, edge servers, and embedded platforms with an accent on quantization, operator substitution, runtime adaptation, and real-device benchmarking. Focus on building repeatable model conversion and deployment pipelines, validating accuracy and latency, and collaborating with embedded engineers and platform vendors on custom kernels and toolchains.

Location: USA — Remote

Base salary: $155,000–$245,000 per year, plus equity, bonus, and benefits.

Company

Deepgram provides real-time speech-to-text, text-to-speech, and voice-agent APIs and deployable voice AI models for developers and organizations.

What you will do

  • Port speech models to non-NVIDIA accelerators, edge servers, and embedded platforms with minimal model changes.
  • Make serving-side decisions involving quantization, precision, operator substitution, graph rewrites, and architecture adjustments.
  • Validate ports on real hardware using accuracy, latency, throughput, and memory benchmarks.
  • Build repeatable model packaging, conversion, versioning, and automated deployment pipelines for edge targets.
  • Collaborate with embedded engineers, platform vendors, and silicon partners on kernels, runtimes, and toolchains.
  • Help address edge deployment concerns such as model security, integrity, automated delivery, and observability.

Requirements

  • Production experience deploying machine learning models to edge or non-NVIDIA hardware; cloud-only or GPU-only serving experience is insufficient.
  • Working knowledge of quantization and precision trade-offs, including INT8, FP16, mixed precision, and calibration.
  • Experience with an edge or vendor inference runtime and conversion toolchain such as ONNX Runtime, TFLite, ExecuTorch, OpenVINO, Qualcomm AI Engine, or a vendor NPU SDK.
  • Ability to read and rewrite model graphs, replace unsupported operators, and adjust architecture parameters while preserving accuracy.
  • Strong Python and PyTorch skills, with production-quality testing, reproducibility, benchmarking, and deployment automation.
  • A builder mindset and ability to scope and deliver ports on unfamiliar platforms.

Nice to have

  • Experience with speech, audio, or streaming and real-time models.
  • Exposure to low-level kernels such as CUDA, Metal, NEON, or vendor DSP code.
  • Experience with model signing, encrypted storage, safe updates, or other deployed-model security measures.
  • Familiarity with Qualcomm, Apple, ARM, Intel, AMD, or custom NPU accelerator families.
  • Experience building internal tooling that improves model porting or deployment speed.

Culture & Benefits

  • AI-first working environment with active use and experimentation with advanced AI tools.
  • Fast-changing work focused on experimentation, adaptation, and continuous learning.
  • Opportunity to ship applied machine learning models to customers across multiple hardware platforms.
  • Equity, annual bonus, and benefits are offered in addition to the base salary.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →