Назад
3 дня назад

Research Engineer (Audio and Speech)

200 000 - 400 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
middle
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Research Engineer (Audio and Speech) (Conversational AI): Building real-time voice-agent models and harnesses for streaming speech, multimodal reasoning, and natural speech generation with an accent on low-latency interaction, full-duplex systems, and production reliability. Focus on designing streaming agent architectures, training audio and speech models, improving recognition and endpointing, and optimizing inference for accuracy, responsiveness, throughput, and cost.

Location: In-office in San Francisco or New York City, United States

Salary: $200K–$400K plus equity

Company

Decagon provides a conversational AI platform for enterprise customer experiences across voice, chat, email, SMS, and other channels.

What you will do

  • Design and build agent harnesses for streaming speech, turn-taking, interruptions, overlapping speech, and continuous interaction.
  • Research and train multimodal and full-duplex models that understand audio, reason, and generate speech.
  • Improve speech recognition, voice activity detection, endpointing, and speech generation across speakers, environments, domains, and languages.
  • Build evaluations using production calls to improve accuracy, latency, naturalness, and task outcomes.
  • Optimize end-to-end inference for responsiveness, throughput, stability, and cost, collaborating with Voice Platform and Infrastructure teams.

Requirements

  • 2+ years of experience in speech, audio ML, multimodal ML, or production machine learning.
  • Experience developing or adapting autoregressive, diffusion, flow-matching, or codec-based speech models.
  • Hands-on experience with streaming agent systems, low-latency inference, production model serving, and evaluation on real-world audio.
  • Fluency in Python and a modern deep-learning framework such as PyTorch, with strong foundations in machine learning and signal processing.
  • Experience taking research ideas from prototype to reliable, measurable production impact.

Nice to have

  • Experience with speech-to-speech or full-duplex models.
  • Experience with telephony, multilingual speech, noisy-channel robustness, speaker adaptation, or expressive speech generation.

Culture & Benefits

  • In-office environment focused on execution, innovation, customer impact, and technical ownership.
  • Medical, dental, vision, life insurance, and disability benefits for full-time employees.
  • Retirement plan, parental leave, and fertility and family-building benefits.
  • Monthly wellness and lifestyle stipend, daily office lunches and snacks, and flexible vacation.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →