Назад
Company hidden
4 дня назад

Product Data Scientist — AI Evaluation & Quality (Remote)

Формат работы
remote (только Europe)/hybrid
Тип работы
fulltime
Грейд
middle
Английский
b2
Страна
Netherlands/Europe
Вакансия из списка Hirify.GlobalВакансия из Hirify RU Global, списка компаний с восточно-европейскими корнями
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Product Data Scientist — AI Evaluation & Quality (AI/LLM): Building offline evaluation suites and online quality dashboards for Finom's AI products with an accent on statistical methodology, regression datasets, judge stability, and production monitoring. Focus on mining real-traffic failure patterns, converting them into regression cases, and translating noisy quality metrics into product decisions.

Location: Remote or hybrid across Europe; listed location: Barcelona

Company

hirify.global is a European fintech startup building an all-in-one B2B financial platform that combines banking, accounting, financial management, and invoicing.

What you will do

  • Own and extend offline evaluation suites across AI products, including capability and regression datasets, judges, and metrics.
  • Build and maintain online quality dashboards covering resolution rate, CSAT, user feedback, LLM-as-judge signals, error rate, and latency.
  • Analyze production traffic to identify failure patterns and convert them into regression cases.
  • Propose improvements to Product and domain experts based on evaluation results.
  • Harden evaluation methodology by addressing judge stability and non-determinism.
  • Translate quality metrics into clear product decisions and trade-offs.

Requirements

  • Professional experience with Python and SQL, including end-to-end analysis.
  • Solid statistical foundation covering sampling, hypothesis testing, variance, and noisy metrics.
  • Analytical mindset focused on business questions and actionable decisions.
  • 3+ years of experience in analyst or data scientist roles, including at least 1 year in a product context.
  • Collaboration with AI engineers, Product, and domain experts.

Nice to have

  • Quality analytics experience for ML systems such as ranking, recommendations, or classification.
  • Hands-on experience evaluating LLM applications, including RAG, agents, tool use, or judges.
  • Experience building LLM agents through side projects, prototypes, or personal experiments.

Culture & Benefits

  • AI-assisted coding is the default authoring environment, with Claude Code used for SQL, Python, analyses, dashboards, and internal scripts.
  • Opportunities for continuous personal and professional development.
  • Stock options available to every team member.
  • Flexibility to travel and work remotely or in a hybrid model across Europe.
  • Access to the Work & Swim program, including one month in a corporate apartment in Cyprus.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →