Mark học AI

Video #49 · How AI gets graded · Part 5/5

How is AI tested for safety and honesty? Red teams and AI judges

Is AI honest? Is it safe? With no right-or-wrong answer key, here's how people grade it.

Watch videoVideo in Vietnamese

Grading what's hard to grade

Last time, you wrote your own test for AI. But some things can't be graded with an answer key: is AI honest, is it safe?

WAY 1

Red teams

The first way: a red team. They play the bad guys and try to trick AI into doing things it shouldn't, so the holes get fixed early.

WAY 2

Refusal tests

The second way: thousands of dangerous requests, to see if AI knows to refuse. And the other way around: it shouldn't refuse normal questions for no reason.

WAY 3

AI grading AI

The third way: use an AI as the judge, grading answers against a written list of criteria. It grades thousands of answers fast, but AI judges make mistakes too, so people need to double-check.

HONESTY

Knowing how to say "I'm not sure"

What about honesty? Testers also ask questions AI can hardly know the answer to. A good AI will say "I'm not sure" instead of making up a confident answer.

RECAP

5 parts in one picture

Looking back at all 5 parts: standard tests, the cramming trap, the voting arena, your own test, and the things that are hard to grade.

NEW SERIES

AI with many senses

AI doesn't just read text anymore: it listens, talks, watches and even makes videos. How? See you in the new series: AI's senses.

This article is based on the video Grading what's hard to grade from the Mark học AI channel. Watch the video (in Vietnamese) to see the animations.