Newer isn't always a better fit
An AI company just released a new model, faster and cheaper. The bank's chatbot team wants to switch right away. Wait. Newer isn't always a better fit.
EVAL SET
The product's own test
They already have a test set, called an eval set: a few hundred real customer questions with the answers they want, plus trick questions they've seen before.
REGRESSION TEST
Old and new take the test
The new model has to redo the whole old test, and its score is compared with the model that's running now. This is called a regression test.
LLM AS A JUDGE
AI grades, people check
Grading hundreds of answers by hand takes a long time, so they use an AI as the judge, called LLM as a judge, grading against clear criteria. Then people double-check the hard ones.
SURPRISE
Better overall, worse where it matters
The results can be surprising: the new model does better on ninety questions, but worse on five about annual fees. Until those five are fixed, no switch.
NOT JUST THE MODEL
Change anything, test again
It's not just switching models. Changing one word in the prompt or adding a document also means testing again. Unlike the benchmarks in the grading series, this eval set belongs to that product alone.
NEXT
When things go wrong, how do you trace it?
The switch is done, and it's live. If something goes wrong, how do you trace it? Part six: logging and tracing.
This article is based on the video Evaluation: test again before switching models from the Mark học AI channel. Watch the video (in Vietnamese) to see the animations.