Mark học AI

Video #46 · How AI gets graded · Part 2/5

What is data contamination? When AI crams for benchmark tests

If AI saw the test before taking it, can you still trust its score?

Watch videoVideo in Vietnamese

What if the test leaks?

Last time, AI took tests called benchmarks. But what if the test leaks ahead of time?

CRAMMING

Studying only the answers

You know the story: someone sees the test early and just memorizes the answers. Sky-high score, but stuck as soon as the test changes.

DATA CONTAMINATION

The test slips into the lessons

AI learns from billions of pages of text online. And AI tests are also posted publicly online. So the test slips into the training data without anyone noticing. This is called data contamination.

FAKE SCORES

Remembering, not understanding

AI has seen the test, so it gets answers right by remembering, not understanding. Change a few numbers in the questions, and the score can drop sharply.

GOODHART'S LAW

When a number becomes the goal

There's a rule named after the economist Goodhart: when a number becomes the goal, it stops measuring well. If everyone races for scores, scores go up, but real skill may not.

NEW TESTS

New tests, hidden tests

That's why people keep writing new tests after AI has finished learning. Some tests are even kept secret and never posted online, so AI can't see them ahead of time.

PART 3

Grading without a test?

So is there a way to grade AI without a test at all? See you in part 3.

This article is based on the video When AI "crams" for the test from the Mark học AI channel. Watch the video (in Vietnamese) to see the animations.