обновлено 28 дней назад
PhD Research Intern (Emotional Speech Generation)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
PhD Research Intern (Emotional Speech Generation) (Deep Learning/Speech and Audio Processing): Developing a controllable multimodal emotional speech generation system in the Dolby Beijing office with an accent on diffusion models, flow matching, and emotion conditioning across text, audio, and image inputs. Focus on designing experiments, analyzing results, transferring research findings into deep learning models, and contributing to academic papers and patents.
Location: Beijing, China; onsite in the Beijing office
Company
Laboratories develops entertainment technologies spanning audio, video, music, movies, gaming, and AR/VR through its Advanced Technology Group.
What you will do
- Design and implement a controllable multimodal emotional speech generation system with research collaborators.
- Investigate generative approaches for speech, including diffusion models and flow matching, to support emotion-aware dialogue editing.
- Build multimodal emotion-conditioning modules using text, audio, and image inputs within a shared emotion embedding space.
- Conduct literature reviews, develop models, design experiments, and analyze results across the full research cycle.
- Contribute to state-of-the-art research, patents, and academic papers.
Requirements
- Working toward a PhD in deep learning for speech and audio processing; PhD candidates are strongly preferred.
- Hands-on experience with emotional or expressive speech generation, including TTS, speech editing, or expressive speech synthesis.
- Practical knowledge of generative models such as diffusion models, flow matching, GANs, or VAEs.
- Experience with speech or multimodal deep learning models, including spoken language models, audio/multimodal language models, or Transformers.
- Strong coding skills with PyTorch and Python.
- Good written communication skills in English.
Nice to have
- Experience with speech disentanglement for speaker identity, linguistic content, and speaking style or emotion.
- Experience with multimodal learning across text, audio, and image.
- Publications in relevant top conferences or journals, including ICASSP, INTERSPEECH, NeurIPS, ICLR, ICML, or ACL.
- Research or project experience in speech emotion transfer, emotional speech synthesis, or dialogue generation.
Culture & Benefits
- Collegial culture with challenging research projects and opportunities to make a visible technical contribution.
- Access to resources across 's audio, video, AR/VR, gaming, music, and movie technology areas.
- Flexible work approach and compensation and benefits package; this role itself is based onsite in Beijing.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →