Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Applied ML Engineer (Edge AI): Porting Deepgram speech models to non-NVIDIA accelerators, edge servers, and embedded platforms with an accent on quantization, operator substitution, runtime adaptation, and real-device benchmarking. Focus on building repeatable model conversion and deployment pipelines, validating accuracy and latency, and collaborating with embedded engineers and platform vendors on custom kernels and toolchains.
Location: USA — Remote
Base salary: $155,000–$245,000 per year, plus equity, bonus, and benefits.
Company
Deepgram provides real-time speech-to-text, text-to-speech, and voice-agent APIs and deployable voice AI models for developers and organizations.
What you will do
- Port speech models to non-NVIDIA accelerators, edge servers, and embedded platforms with minimal model changes.
- Make serving-side decisions involving quantization, precision, operator substitution, graph rewrites, and architecture adjustments.
- Validate ports on real hardware using accuracy, latency, throughput, and memory benchmarks.
- Build repeatable model packaging, conversion, versioning, and automated deployment pipelines for edge targets.
- Collaborate with embedded engineers, platform vendors, and silicon partners on kernels, runtimes, and toolchains.
- Help address edge deployment concerns such as model security, integrity, automated delivery, and observability.
Requirements
- Production experience deploying machine learning models to edge or non-NVIDIA hardware; cloud-only or GPU-only serving experience is insufficient.
- Working knowledge of quantization and precision trade-offs, including INT8, FP16, mixed precision, and calibration.
- Experience with an edge or vendor inference runtime and conversion toolchain such as ONNX Runtime, TFLite, ExecuTorch, OpenVINO, Qualcomm AI Engine, or a vendor NPU SDK.
- Ability to read and rewrite model graphs, replace unsupported operators, and adjust architecture parameters while preserving accuracy.
- Strong Python and PyTorch skills, with production-quality testing, reproducibility, benchmarking, and deployment automation.
- A builder mindset and ability to scope and deliver ports on unfamiliar platforms.
Nice to have
- Experience with speech, audio, or streaming and real-time models.
- Exposure to low-level kernels such as CUDA, Metal, NEON, or vendor DSP code.
- Experience with model signing, encrypted storage, safe updates, or other deployed-model security measures.
- Familiarity with Qualcomm, Apple, ARM, Intel, AMD, or custom NPU accelerator families.
- Experience building internal tooling that improves model porting or deployment speed.
Culture & Benefits
- AI-first working environment with active use and experimentation with advanced AI tools.
- Fast-changing work focused on experimentation, adaptation, and continuous learning.
- Opportunity to ship applied machine learning models to customers across multiple hardware platforms.
- Equity, annual bonus, and benefits are offered in addition to the base salary.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
4 дня назад
AI Research Engineer
100 000 - 150 000$
4 дня назад
Research Engineer (AI)
190 000 - 240 000$
Baseten
16 часов назад
Software Engineer (AI)
180 000 - 360 000$
Anthropic
16 часов назад
Performance Engineer (AI)
280 000 - 850 000$
Baseten
16 часов назад
Gpu Kernel Engineer (AI)
180 000 - 360 000$
4 дня назад
Staff ML Engineer (AI)
130 000 - 190 000$