Назад
Company hidden
7 дней назад

Research Scientist, Speech & Audio (AI)

160 000 - 185 000$
Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Research Scientist, Speech & Audio (AI) (ASR/TTS and Audio Models): Designing and evaluating audio data, benchmarks, and speech-model experiments with an accent on multilingual coverage, transcription quality, robustness, and streaming performance. Focus on building adversarial evaluations, fine-tuning models, measuring data-driven improvements, and publishing reproducible research.

Location: Remote - United States

Salary: $160,000–$185,000 per year, based on experience, skills, and qualifications.

Company

hirify.global is a global data engineering company that provides data, evaluation frameworks, and human expertise for generative AI systems.

What you will do

  • Define data structures, annotation schemas, sampling strategies, and evaluation criteria for ASR, text-to-speech, speech-to-speech, conversational voice, diarization, verification, audio-language, and streaming models.
  • Build evaluation methodologies covering semantic accuracy, accent and noise robustness, code-switching, diarization error rate, speech naturalness, intelligibility, and streaming latency.
  • Design audio coverage across languages, accents, dialects, acoustic conditions, speaker demographics, emotional and paralinguistic characteristics, and single- or multi-speaker settings.
  • Partner with transcription, linguistics, audio solutions, engineering, annotation, and subject-matter teams to turn model objectives into operational data and evaluation programs.
  • Fine-tune and evaluate models on hirify.global data, using ablations to connect data decisions with measurable improvements.
  • Design adversarial evaluations, document findings, and publish benchmarks, methodologies, and research papers.

Requirements

  • Approximately 5+ years of hands-on industry experience in speech or audio machine learning; a current PhD research agenda may offset experience at the lower end.
  • Bachelor’s degree in computer science, electrical engineering, or a related technical or quantitative field.
  • Hands-on experience training and evaluating ASR, TTS, speech-to-speech, speaker, or audio-language models, with strong PyTorch fundamentals.
  • Experience with ESPnet, NeMo, SpeechBrain, or Kaldi, Hugging Face, forced alignment, WER/CER, and advanced speech evaluation metrics.
  • Experience with multilingual, accented, dialectal, low-resource, or code-switched speech and synthetic or augmented audio.
  • First-author publications or strong open-source contributions at venues such as Interspeech, ICASSP, ASRU, SLT, or NeurIPS.

Nice to have

  • Advanced degree in a relevant field.
  • Experience with responsible-AI evaluation and red-teaming, including spoofing, voice-cloning robustness, or bias across accents and languages.

Culture & Benefits

  • Direct collaboration with customers and frontier AI laboratories developing speech and audio models.
  • Work spans research, data design, evaluation methodology, annotation, and engineering implementation.
  • Practical experience is valued alongside formal credentials.
  • Research findings can be developed into benchmarks, methodologies, open-source contributions, and papers.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →