Назад
Company hidden
7 часов назад

Machine Learning Researcher (Speech AI)

140 000 - 250 000$
Формат работы
remote (только USA)/onsite
Тип работы
fulltime
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Machine Learning Researcher (Speech AI): Building and scaling TTS, speech-to-text, neural audio codec, and real-time voice systems with an accent on expressive speech generation, robust transcription, and efficient large-scale training. Focus on designing novel models, running rigorous ablation studies, and deploying low-latency inference for enterprise voice agents.

Location: San Francisco, CA or Remote within the United States

Salary: $140,000–$250,000 annually, with equity and bonus opportunities. The posting also mentions a competitive salary range of $160,000–$250,000.

Company

hirify.global builds AI phone agents and the models and infrastructure that enable natural, reliable voice interactions for enterprises at scale.

What you will do

  • Design and train expressive, controllable text-to-speech models and neural audio codec-based architectures.
  • Build and fine-tune speech-to-text systems that handle accents, noise, telephony artifacts, code switching, and conversational nuance.
  • Research neural audio codecs with high compression, low perceptual loss, and support for controllable speech generation.
  • Curate large multilingual audio datasets and develop staged training, filtering, and distributed GPU pipelines.
  • Optimize large-model training and real-time inference for latency, throughput, cost, memory efficiency, and reliability.
  • Design ablation studies and evaluate research through objective metrics, perceptual testing, offline benchmarks, and production metrics.

Requirements

  • Experience with self-supervised, multimodal, or generative modeling and the ability to derive and implement new formulations.
  • Hands-on experience building or scaling TTS, STT, or neural audio codec systems.
  • Experience with large-scale distributed training and serving models on modern accelerators.
  • Knowledge of quantization, kernel optimization, memory efficiency, and real-time or streaming speech constraints.
  • Strong experimental skills, including controlled experiments, ablation studies, and meaningful evaluation.
  • Research publications or open-source contributions in speech or language AI, or a PhD in ML, AI, or a related field, are expected as evidence of research impact.

Nice to have

  • Background in real-time speech systems or telephony.
  • Experience with large-scale speech datasets and real-world audio variability.
  • Strong intuition for audio quality, prosody, emotional expression, and conversational dynamics.

Culture & Benefits

  • Fast-moving startup environment with ownership from research through production deployment.
  • Healthcare, dental, and vision benefits.
  • Meaningful equity in a fast-growing company.
  • Access to the tools needed to support research and development.
  • San Francisco office at Levi's Plaza with rooftop views.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →