Назад
Company hidden
7 часов назад

Machine Learning Scientist (Speech AI)

Формат работы
remote (только USA)
Тип работы
fulltime
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Machine Learning Scientist (Speech AI): Designing, training, and evaluating speech synthesis and speech understanding models for production voice AI with an accent on full-duplex multimodal architectures, neural speech representations, and prosodic control. Focus on building rigorous objective and perceptual evaluations, improving data quality, and developing scalable PyTorch training pipelines.

Location: Remote within the United States

Company

hirify.global builds voice AI and text-to-speech models for high-volume enterprise customer experiences, using a proprietary studio-quality conversational speech corpus.

What you will do

  • Design, train, and evaluate autoregressive and non-autoregressive speech synthesis models.
  • Research full-duplex and half-duplex multimodal architectures, including unified speech-to-speech systems.
  • Develop and iterate on speech representations such as neural codecs, semantic tokens, mel features, and continuous latents.
  • Build objective and perceptual evaluations focused on speech quality and prosodic control.
  • Collaborate with linguists on text-to-speech frontend behavior, including pronunciation and prosody.
  • Build data and training pipelines and help move research models toward production.

Requirements

  • Deep familiarity with speech synthesis research, including Tacotron, FastSpeech, VITS, VALL-E, and codec-language-model approaches.
  • Hands-on experience training neural codecs such as EnCodec, DAC, or Mimi and working with multiple speech representation choices.
  • Experience with full- or half-duplex multimodal modeling, streaming speech-to-speech systems, or related architectures.
  • Strong attention to data quality, annotation pipelines, and evaluation-set integrity.
  • Working knowledge of TTS frontend technologies, including grapheme-to-phoneme conversion, normalization, and prosody.
  • Strong PyTorch fundamentals, including training loops, distributed training, and model internals; PhD or equivalent research experience in speech, audio, machine learning, or computational linguistics is expected, unless offset by a strong research track record.

Nice to have

  • Multilingual TTS experience.
  • Background in prosody or paralinguistics.
  • Published work in speech, audio, or core machine learning venues.
  • Experience with quantization, distillation, or streaming inference in production.

Culture & Benefits

  • Remote-friendly work environment.
  • Visa sponsorship available.
  • Access to a proprietary full-duplex, studio-quality conversational speech corpus.
  • Compute and tooling for research and development.
  • Competitive base compensation with meaningful early-stage equity.
  • High ownership, strong standards, direct founder collaboration, and low bureaucracy.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →