7 часов назад
Machine Learning Scientist (Speech AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Machine Learning Scientist (Speech AI): Designing, training, and evaluating speech synthesis and speech understanding models for production voice AI with an accent on full-duplex multimodal architectures, neural speech representations, and prosodic control. Focus on building rigorous objective and perceptual evaluations, improving data quality, and developing scalable PyTorch training pipelines.
Location: Remote within the United States
Company
builds voice AI and text-to-speech models for high-volume enterprise customer experiences, using a proprietary studio-quality conversational speech corpus.
What you will do
- Design, train, and evaluate autoregressive and non-autoregressive speech synthesis models.
- Research full-duplex and half-duplex multimodal architectures, including unified speech-to-speech systems.
- Develop and iterate on speech representations such as neural codecs, semantic tokens, mel features, and continuous latents.
- Build objective and perceptual evaluations focused on speech quality and prosodic control.
- Collaborate with linguists on text-to-speech frontend behavior, including pronunciation and prosody.
- Build data and training pipelines and help move research models toward production.
Requirements
- Deep familiarity with speech synthesis research, including Tacotron, FastSpeech, VITS, VALL-E, and codec-language-model approaches.
- Hands-on experience training neural codecs such as EnCodec, DAC, or Mimi and working with multiple speech representation choices.
- Experience with full- or half-duplex multimodal modeling, streaming speech-to-speech systems, or related architectures.
- Strong attention to data quality, annotation pipelines, and evaluation-set integrity.
- Working knowledge of TTS frontend technologies, including grapheme-to-phoneme conversion, normalization, and prosody.
- Strong PyTorch fundamentals, including training loops, distributed training, and model internals; PhD or equivalent research experience in speech, audio, machine learning, or computational linguistics is expected, unless offset by a strong research track record.
Nice to have
- Multilingual TTS experience.
- Background in prosody or paralinguistics.
- Published work in speech, audio, or core machine learning venues.
- Experience with quantization, distillation, or streaming inference in production.
Culture & Benefits
- Remote-friendly work environment.
- Visa sponsorship available.
- Access to a proprietary full-duplex, studio-quality conversational speech corpus.
- Compute and tooling for research and development.
- Competitive base compensation with meaningful early-stage equity.
- High ownership, strong standards, direct founder collaboration, and low bureaucracy.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
1 день назад
Senior Machine Learning Engineer (AI)
225 000 - 300 000$
7 часов назад
ML Data Engineer
Sber
1 день назад
Senior Deep Learning Engineer (Speech/Audio)
6 часов назад
ML Data Engineer (AI)
Sber
1 день назад
Senior Deep Learning Engineer (Speech LLM)
500 000₽
6 часов назад