Leaderboards aren't enough
A leaderboard says this model is the best. But whether your product is good, a leaderboard can't tell you.
WAY 1
A set of sample questions
The first way: a set of sample questions, taken from real customer questions, with the right answers. A bakery might have fifty: cake prices, opening hours, how to order.
WAY 1
Rerun it every time you change something
Every time you change something, the model or the instructions, rerun the whole set. Any question that used to be right and is now wrong shows up right away, before customers see it.
WAY 2
Signals from real users
The second way: listen to real users. Not just the like and dislike buttons. If they ask the same thing again, they weren't satisfied. If they edit the answer before using it, it was close. If they leave halfway, something's wrong.
WAY 3
A/B testing
The third way: an A/B test. Half the customers see the old version, half see the new one. Whichever gets more cake orders wins.
RECAP
Three ways together
All three together: the question set keeps quality steady, real signals show where it hurts, and A/B tests help you choose a direction.
NEXT
Every answer costs money
But every answer being measured costs money. How much, and who pays? See you in part 5.
This article is based on the video Measuring an AI product from the Mark học AI channel. Watch the video (in Vietnamese) to see the animations.