Mark học AI

Video #83 · Behind the scenes of an AI system · Part 5/7

What is evaluation? Why you test again before switching AI models

An AI company releases a new model, faster and cheaper. Switch right away? Wait: newer isn't always a better fit.

Follow on YouTubeComing soonVideo in Vietnamese

Newer isn't always a better fit

An AI company just released a new model, faster and cheaper. The bank's chatbot team wants to switch right away. Wait. Newer isn't always a better fit.

EVAL SET

The product's own test

They already have a test set, called an eval set: a few hundred real customer questions with the answers they want, plus trick questions they've seen before.

REGRESSION TEST

Old and new take the test

The new model has to redo the whole old test, and its score is compared with the model that's running now. This is called a regression test.

LLM AS A JUDGE

AI grades, people check

Grading hundreds of answers by hand takes a long time, so they use an AI as the judge, called LLM as a judge, grading against clear criteria. Then people double-check the hard ones.

SURPRISE

Better overall, worse where it matters

The results can be surprising: the new model does better on ninety questions, but worse on five about annual fees. Until those five are fixed, no switch.

NOT JUST THE MODEL

Change anything, test again

It's not just switching models. Changing one word in the prompt or adding a document also means testing again. Unlike the benchmarks in the grading series, this eval set belongs to that product alone.

NEXT

When things go wrong, how do you trace it?

The switch is done, and it's live. If something goes wrong, how do you trace it? Part six: logging and tracing.

This article is based on the video Evaluation: test again before switching models from the Mark học AI channel. Watch the video (in Vietnamese) to see the animations.