Назад
Company hidden
5 дней назад

PhD Research Intern - Emotional Speech Generation

Формат работы
onsite
Тип работы
fulltime
Грейд
trainee
Английский
b2
Страна
China
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
PhD Research Intern - Emotional Speech Generation (Deep Learning/Speech): Designing and implementing controllable multimodal emotional speech generation systems with an accent on generative modeling, emotion conditioning, and speech-audio processing. Focus on building shared emotion representations across text, audio, and image inputs, developing diffusion and flow-matching models, and conducting research toward state-of-the-art results.

Location: Onsite in the Beijing office, China

Company

Entertainment technology organization developing innovations across audio, video, music, gaming, and immersive media.

What you will do

  • Design and implement controllable multimodal emotional speech generation systems.
  • Research and apply diffusion models, flow matching, and related generative approaches for emotion-aware dialogue editing.
  • Build emotion conditioning modules using text, reference speech, and facial-expression inputs.
  • Align multimodal inputs into a shared emotion embedding space.
  • Conduct literature reviews, experiments, analysis, and hands-on model development with research collaborators.
  • Contribute to patents and academic publications while pursuing state-of-the-art results.

Requirements

  • PhD candidate status is strongly preferred in deep learning for speech and audio processing.
  • Hands-on experience with emotional or expressive speech generation, TTS, speech editing, or expressive speech synthesis.
  • Practical understanding of generative models such as diffusion models, flow matching, GANs, or VAEs.
  • Experience with Transformer-based, spoken-language, audio, or multimodal language models.
  • Strong programming skills in Python and PyTorch.
  • Good written communication skills in English and the ability to analyze and communicate complex information.

Nice to have

  • Experience with speech disentanglement for speaker identity, linguistic content, and speaking style or emotion.
  • Multimodal learning experience involving text, audio, and image modalities.
  • Publications in relevant conferences or journals such as ICASSP, INTERSPEECH, NeurIPS, ICLR, ICML, or ACL.
  • Research experience in speech emotion transfer, emotional speech synthesis, or dialogue generation.

Culture & Benefits

  • Collaborative culture with challenging research and technology projects.
  • Access to resources across audio, video, AR/VR, gaming, music, and movie technologies.
  • Flexible work approach designed to support different ways of working, with this role based onsite in Beijing.
  • Compensation and benefits are offered for the position.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →