3 дня назад
Senior Data Scientist (AI Evaluation & Quality)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Data Scientist (AI Evaluation & Quality) (AI/Data Science): Building evaluation systems and quality monitoring for AI products, including a financial co-pilot, voice agent, and internal AI-powered processes, with an accent on offline evaluation suites, online quality dashboards, and production feedback loops. Focus on analyzing failure patterns, creating regression cases, stabilizing LLM judges, and turning noisy metrics into product decisions.
Location: Vilnius; remote or hybrid work across Europe.
Company
is a European fintech startup building an all-in-one B2B financial platform for entrepreneurs, combining banking, accounting, financial management, and invoicing.
What you will do
- Own and extend offline evaluation suites across AI products, including datasets, judges, metrics, capability tests, and regression tests.
- Build and maintain online quality dashboards covering resolution rate, CSAT, user feedback, LLM-as-judge signals, error rate, and latency.
- Analyze real-traffic failures, identify recurring patterns, and convert findings into regression cases.
- Collaborate with AI engineers, Product, and domain experts to propose product and quality improvements.
- Improve evaluation methodology by addressing judge stability and non-deterministic results.
- Translate analytical findings into clear trade-offs and product decisions during weekly syncs.
Requirements
- 3+ years of experience in analyst or data scientist roles, including at least one year in a product context.
- Strong Python and SQL skills for end-to-end analysis.
- Solid statistical foundation, including sampling, hypothesis testing, variance, and noisy metrics.
- Analytical approach focused on business questions and measurable decisions.
- Fluency with AI-assisted coding or genuine readiness to become fluent quickly.
- Ability to work remotely or in a hybrid model across Europe.
Nice to have
- Quality analytics experience for ML systems such as ranking, recommendations, or classification.
- Hands-on experience evaluating LLM applications, including RAG, agents, tool use, or judges.
- Experience building LLM agents through side projects or personal experiments.
Culture & Benefits
- AI-assisted coding is the default authoring environment, with Claude Code used for SQL, Python, analyses, dashboards, and internal scripts.
- Opportunities for continuous personal and professional development.
- Stock options available to every team member.
- Supportive, friendly, and eco-conscious corporate culture.
- Work & Swim program offering one month in a corporate apartment in Cyprus.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
16 часов назад
Senior Data Science Consultant (AI)
3 дня назад
Model Risk Specialist (Credit Risk/AI)
Andersen Lab
8 часов назад
Senior Data Scientist (Warsaw, Porto, Amsterdam)
3 000 - 4 500€
4 дня назад
Software Engineer (Enterprise AI)
144 500 - 180 600$
8 часов назад
Senior AI/Machine Learning Engineer (LLM)
4 дня назад