Назад
2 дня назад

Senior Software Engineer - Model Evaluation & AI Systems (AI)

180 000 - 240 000$
Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/

TL;DR

Senior Software Engineer - Model Evaluation & AI Systems (AI): Building and maintaining automated evaluation pipelines and infrastructure for speech, audio, and multimodal AI models with an accent on correctness, reproducibility, and scalability. Focus on creating evaluation methodologies, designing pass/fail gates, and developing continuous monitoring systems to detect quality regressions in production.

Location: Remote (USA)

Salary: $180,000 – $240,000 base + Equity + 10% Annual Bonus

Company

Deepgram is a leading Voice AI platform providing real-time APIs for speech-to-text, text-to-speech, and production-grade voice agents.

What you will do

  • Define and build evaluation methodologies for STT, TTS, LLM, RAG, agent, and multimodal systems.
  • Design and maintain automated evaluation pipelines focusing on WER, hallucination detection, latency, and time-to-first-byte.
  • Build scalable evaluation infrastructure, including harnesses and result-aggregation pipelines for production models and GPU clusters.
  • Translate research benchmarks into automated, enforceable pass/fail gates.
  • Operate canaries and continuous-monitoring systems to detect quality regressions.
  • Integrate evaluation and quality gates into CI/CD pipelines to ensure continuous verification.

Requirements

  • BS, MS, or PhD in Computer Science, AI, Applied Math, or equivalent experience.
  • 5+ years of professional software or QA engineering experience with a track record of shipping test infrastructure.
  • Solid backend experience in Python, Rust, Go, or similar languages.
  • Experience designing automated test pipelines, evaluation frameworks, or data-processing systems.
  • Strong analytical skills and ability to reason about metrics and statistical variation.
  • Must be based in the USA

Nice to have

  • Hands-on experience evaluating LLMs, RAG pipelines, agents, or multimodal models.
  • Experience with React Native or cross-platform mobile frameworks for internal tooling.
  • Familiarity with voice/audio metrics such as WER, MOS, or TTFB.
  • Prior contributions to open-source projects.
  • Experience with cloud infrastructure, containerized environments, and monitoring tools like Grafana.

Culture & Benefits

  • AI-first culture: active use and experimentation with advanced AI tools is a core requirement.
  • High-paced environment with rapid evolution of day-to-day tasks.
  • Competitive compensation package including base salary, equity, and annual bonus.
  • Opportunity to work with state-of-the-art voice-native foundation models.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →