5 дней назад
PhD Research Intern - Emotional Speech Generation
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
PhD Research Intern - Emotional Speech Generation (Deep Learning/Speech): Designing and implementing controllable multimodal emotional speech generation systems with an accent on generative modeling, emotion conditioning, and speech-audio processing. Focus on building shared emotion representations across text, audio, and image inputs, developing diffusion and flow-matching models, and conducting research toward state-of-the-art results.
Location: Onsite in the Beijing office, China
Company
Entertainment technology organization developing innovations across audio, video, music, gaming, and immersive media.
What you will do
- Design and implement controllable multimodal emotional speech generation systems.
- Research and apply diffusion models, flow matching, and related generative approaches for emotion-aware dialogue editing.
- Build emotion conditioning modules using text, reference speech, and facial-expression inputs.
- Align multimodal inputs into a shared emotion embedding space.
- Conduct literature reviews, experiments, analysis, and hands-on model development with research collaborators.
- Contribute to patents and academic publications while pursuing state-of-the-art results.
Requirements
- PhD candidate status is strongly preferred in deep learning for speech and audio processing.
- Hands-on experience with emotional or expressive speech generation, TTS, speech editing, or expressive speech synthesis.
- Practical understanding of generative models such as diffusion models, flow matching, GANs, or VAEs.
- Experience with Transformer-based, spoken-language, audio, or multimodal language models.
- Strong programming skills in Python and PyTorch.
- Good written communication skills in English and the ability to analyze and communicate complex information.
Nice to have
- Experience with speech disentanglement for speaker identity, linguistic content, and speaking style or emotion.
- Multimodal learning experience involving text, audio, and image modalities.
- Publications in relevant conferences or journals such as ICASSP, INTERSPEECH, NeurIPS, ICLR, ICML, or ACL.
- Research experience in speech emotion transfer, emotional speech synthesis, or dialogue generation.
Culture & Benefits
- Collaborative culture with challenging research and technology projects.
- Access to resources across audio, video, AR/VR, gaming, music, and movie technologies.
- Flexible work approach designed to support different ways of working, with this role based onsite in Beijing.
- Compensation and benefits are offered for the position.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
Jump Trading
5 дней назад
Campus AI ML Researcher (Fintech)
3 дня назад
Machine Learning Engineer
3 дня назад
Machine Learning Engineer (Machine Learning)
3 дня назад
Senior Quantitative Researcher (Machine Learning)
120 000 - 200 000$
2 дня назад
Data Scientist (Gaming)
140 000 - 165 000$
3 дня назад