4 дня назад
Product Data Scientist — AI Evaluation & Quality (Remote)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Product Data Scientist — AI Evaluation & Quality (AI/LLM): Building offline evaluation suites and online quality dashboards for Finom's AI products with an accent on statistical methodology, regression datasets, judge stability, and production monitoring. Focus on mining real-traffic failure patterns, converting them into regression cases, and translating noisy quality metrics into product decisions.
Location: Remote or hybrid across Europe; listed location: Barcelona
Company
is a European fintech startup building an all-in-one B2B financial platform that combines banking, accounting, financial management, and invoicing.
What you will do
- Own and extend offline evaluation suites across AI products, including capability and regression datasets, judges, and metrics.
- Build and maintain online quality dashboards covering resolution rate, CSAT, user feedback, LLM-as-judge signals, error rate, and latency.
- Analyze production traffic to identify failure patterns and convert them into regression cases.
- Propose improvements to Product and domain experts based on evaluation results.
- Harden evaluation methodology by addressing judge stability and non-determinism.
- Translate quality metrics into clear product decisions and trade-offs.
Requirements
- Professional experience with Python and SQL, including end-to-end analysis.
- Solid statistical foundation covering sampling, hypothesis testing, variance, and noisy metrics.
- Analytical mindset focused on business questions and actionable decisions.
- 3+ years of experience in analyst or data scientist roles, including at least 1 year in a product context.
- Collaboration with AI engineers, Product, and domain experts.
Nice to have
- Quality analytics experience for ML systems such as ranking, recommendations, or classification.
- Hands-on experience evaluating LLM applications, including RAG, agents, tool use, or judges.
- Experience building LLM agents through side projects, prototypes, or personal experiments.
Culture & Benefits
- AI-assisted coding is the default authoring environment, with Claude Code used for SQL, Python, analyses, dashboards, and internal scripts.
- Opportunities for continuous personal and professional development.
- Stock options available to every team member.
- Flexibility to travel and work remotely or in a hybrid model across Europe.
- Access to the Work & Swim program, including one month in a corporate apartment in Cyprus.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
3 дня назад
Model Risk Specialist (Credit Risk/AI)
16 часов назад
Senior Data Science Consultant (AI)
Andersen Lab
8 часов назад
Senior Data Scientist (Warsaw, Porto, Amsterdam)
3 000 - 4 500€
3 дня назад
Data Scientist (NLP/Search)
Andersen Lab
4 дня назад
Data Scientist / Data Analyst
4 дня назад