Назад
обновлено 5 дней назад

Director of Research (Text-to-Speech)

213 000 - 266 300$
Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
director
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Director of Research (Text-to-Speech): Owning the end-to-end TTS research program, from research strategy and neural audio modeling to production-grade speech-generation models, with an accent on prosody, expressiveness, controllability, multilingual speech, and voice consistency. Focus on designing experiments, building evaluation systems, making high-impact technical bets, and developing research teams through technical leaders.

Location: Remote within the United States; listed locations include Ann Arbor, Michigan and San Francisco, California.

Base salary: $213,000–$266,300 in most U.S. locations; $262,600–$328,300 in San Francisco, New York City, and Seattle. Equity and bonus are also offered.

Company

Deepgram provides real-time voice AI APIs and self-hosted software for speech-to-text, text-to-speech, and production voice agents.

What you will do

  • Own the text-to-speech research and model roadmap from technical strategy through production delivery.
  • Advance neural audio modeling, prosody, expressiveness, controllability, multilingual speech, voice identity, training, post-training, and inference performance.
  • Review research, design experiments, diagnose model failures, challenge assumptions, and solve high-leverage technical problems.
  • Build automated and human evaluation systems that explain why generative speech models improve.
  • Lead individual contributors and technical lead managers while hiring, developing, and setting direction across research sub-teams.
  • Partner with engineering and product leadership on ship-readiness and represent TTS research internally and externally.

Requirements

  • Deep expertise in modern TTS, speech generation, or audio generative modeling, including personally training and improving large-scale neural models.
  • Strong command of speech-generation challenges involving naturalness, expressiveness, controllability, robustness, voice consistency, and inference cost.
  • Experience setting research direction, prioritizing experiments, allocating compute and researcher time, and discontinuing ineffective approaches.
  • Experience leading researchers and research engineers through other technical leaders while remaining technically influential.
  • Experience deploying TTS or generative-audio models at meaningful production scale.
  • AI must be a default mode of work, with active use and experimentation with advanced AI tools.

Nice to have

  • Experience building or substantially scaling a high-performing AI research organization.
  • Sophisticated evaluation systems for generative speech, expressive or multilingual generation, voice cloning and adaptation, or controllable generation.
  • Publications, open-source contributions, patents, or invited talks in speech synthesis, neural audio codecs, speech language models, or multimodal models.
  • Experience in fast-moving startup or research environments taking models from idea to production.

Culture & Benefits

  • AI-first operating model with continuous experimentation and rapid adaptation.
  • Remote work within the United States.
  • Equity and performance bonus in addition to base salary.
  • Work focused on frontier voice AI research and production-scale model deployment.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →