Назад
Company hidden
3 дня назад

Principal Speech Data Linguist (AI)

160 000 - 185 000$
Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Principal Speech Data Linguist (AI): Defining linguistic standards, quality frameworks, and human-in-the-loop workflows for high-volume speech segmentation and transcription with an accent on IPA, acoustic analysis, diarization, multilingual data, and difficult speech conditions. Focus on designing annotation specifications, adjudicating complex cases, measuring quality, and connecting transcription choices to ASR, TTS, and speech-understanding model performance.

Location: Remote - United States

Salary: $160,000–$185,000 per year, based on experience, skills, and qualifications.

Company

hirify.global is a global data engineering company providing data, evaluation frameworks, platforms, and human expertise for generative AI builders and adopters.

What you will do

  • Define transcription and segmentation standards, style guides, annotation conventions, timestamping rules, speaker labels, diarization labels, and non-speech event conventions.
  • Build quality frameworks covering rubrics, error taxonomies, adjudication, inter-annotator agreement, and human QA at scale.
  • Own the end-to-end quality lifecycle for speech data, including preprocessing, acceptance checks, post-processing, validation, reporting, and delivery packaging.
  • Design human-in-the-loop workflows that focus expert review on cases ASR models cannot reliably handle.
  • Address accented, dialectal, multilingual, low-resource, overlapping, domain-specific, and noisy speech.
  • Partner with research scientists, delivery operations, customers, and frontier labs to translate model objectives into speech-data specifications and best practices.

Requirements

  • Typically 8+ years of industry experience in transcription, segmentation, and speech-data quality, including authorship of standards.
  • Bachelor’s degree in linguistics, phonetics, computational linguistics, or a closely related field; an advanced degree is preferred.
  • Strong knowledge of phonetics, phonology, sociolinguistics, IPA, prosody, disfluencies, dialects, and registers.
  • Hands-on experience with acoustic and phonetic analysis, spectrograms, formants, pitch, prosody, segment boundaries, and tools such as Praat.
  • Experience with Whisper, commercial ASR engines, forced alignment, ELAN, audio segmentation, diarization, Python, regular expressions, WER, and inter-annotator agreement.
  • Multilingual experience with accented, dialectal, and code-switched speech, plus strong written and verbal communication skills.

Nice to have

  • Low-resource language experience.
  • Advanced degree in a relevant field.
  • Knowledge of responsible-AI topics in speech, including accent and dialect bias, privacy, and consent in voice data.

Culture & Benefits

  • Remote work within the United States.
  • Work on speech-data workflows supporting generative AI builders, adopters, and frontier labs.
  • Collaboration with research scientists, delivery operations, customers, and expert reviewers.
  • Opportunity to shape standards and methodologies as speech models and human-in-the-loop workflows evolve.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →