Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
TL;DR
Staff Research Engineer (AI): Developing next-generation voice-to-voice and multimodal interactive models with an accent on low-latency synthesis and natural conversational flow. Focus on designing multi-modal system architectures and implementing end-to-end generative pipelines from pretraining to production.
Location: Remote (Europe)
Company
Synthesia is the world’s leading AI video platform for business, helping enterprises enhance visual communication through human-interactive models.
What you will do
- Shape the roadmap for new model capabilities and propose novel multi-modal system architectures (text and voice).
- Develop and evaluate streaming and conversational systems for low-latency, interactive voice-video synthesis.
- Implement generative pipelines from pretraining through post-training, utilizing DPO, fine-tuning, and distillation.
- Integrate novel architectures including neural codecs, diffusion, and flow-matching to enhance realism.
- Define new evaluation metrics for conversational systems and curate specialized datasets.
- Ship optimized models to production and iterate based on direct customer feedback.
Requirements
- Strong understanding of generative modeling applied to sequential or multimodal data.
- Hands-on experience with LLMs or transformer-based architectures.
- High proficiency in PyTorch, including distributed training and model optimization.
- Solid grasp of time-series modeling and tokenization in audio, speech, or video contexts.
- Proven experience training deep learning models end-to-end, from data preparation to evaluation.
- Must be based in Europe.
Nice to have
- Experience shipping generative models into live products used at a meaningful scale.
- Background in interactive systems where latency and responsiveness were primary constraints.
- Experience with large-scale LLM training resulting in strong reasoning capabilities.
- Original research contributions in top-tier venues such as NeurIPS, CVPR, ICML, ICLR, or Interspeech.
- Familiarity with SOTA audio/speech architectures like flow-matching or autoregressive decoders.
Culture & Benefits
- Opportunity to work at a high-growth AI company with a $4 billion valuation and premier VC backing.
- Collaborate within a 40+ person R&D department on cutting-edge interactive AI.
- High level of ownership over critical components and the ability to drive broad research directions.
- Remote-first work environment across Europe.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →