Назад
11 дней назад

Senior Engineer (Voice AI)

Формат работы
remote (Global)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
France/US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Engineer (Voice AI): Building Hugging Face’s open voice-agent stack, including the speech-to-speech library and hf-voice product, with an accent on realtime inference, developer APIs, and open-source infrastructure. Focus on designing streaming protocols, serving voice models on GPUs, managing concurrency and autoscaling, and taking the platform from demo to production reliability.

Location: Remote

Company

Hugging Face builds an open platform and open-source tools for AI developers, including models, datasets, applications, and machine learning infrastructure.

What you will do

  • Own the architecture of the open-source speech-to-speech library, including pipeline design, latency budgets, and realtime-loop reliability.
  • Integrate ASR, TTS, and end-to-end speech models while maintaining clean abstractions.
  • Design hf-voice developer APIs and streaming protocols covering session lifecycle, WebSockets/WebRTC transport, authentication, errors, and versioning.
  • Build realtime GPU inference serving with concurrency, autoscaling, observability, and cost optimization.
  • Take the product from demo to production through load testing, SLOs, and graceful degradation.
  • Write documentation and examples, support Reachy Mini deployments, and collaborate with the Hub, inference, and open-source communities.

Requirements

  • Senior engineering experience with autonomous ownership of substantial architecture.
  • Experience building developer-facing infrastructure for AI or developer tools, such as inference APIs or agent infrastructure.
  • Substantial open-source contributions to a Python library and comfort with async Python and distributed systems.
  • Experience shipping realtime systems involving streaming, WebSockets, WebRTC, audio, video, or live inference.
  • Production experience with LLMs or multimodal models, plus clear written communication and asynchronous collaboration.
  • Strong motivation to work on voice and conversational AI.

Nice to have

  • Contributions to voice-agent frameworks or low-level inference runtimes.
  • Experience with ASR, TTS, speech-model evaluation, GPU serving, quantization, or on-device inference.
  • Audio pipeline knowledge, including VAD, echo cancellation, jitter buffers, barge-in, and turn detection.
  • Experience shipping to embedded or robotics targets and a public record of talks, blog posts, or demos.

Culture & Benefits

  • Remote options, flexible working hours, and a distributed workplace with offices in NYC and Paris.
  • Health, dental, and vision benefits for employees and dependents.
  • Parental leave, flexible paid time off, and workstation equipment when needed.
  • Reimbursement for relevant conferences, training, and education.
  • Company equity for employees and support for diversity, equity, and inclusion.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →