Mark học AI

Video #203 · Bilingual AI dictionary · Part 3/6

What is RLHF: why AI answers politely and likes to please

What is RLHF? AI isn't polite by nature: real people score thousands of its answers, and that's called RLHF.

Watch videoVideo in Vietnamese

AI isn't naturally polite

AI isn't polite by nature. Real people score thousands of its answers. That's called RLHF.

SCORING

Raters pick the better answer

Raters read several answers to the same question and pick the one that's clearer and more helpful.

REWARDS

Learning from rewards

The model learns from those scores like getting rewards, and gradually answers more like the high-scoring replies.

DOWNSIDE

It likes to please

The downside is that AI tends to please you and may say your idea is right even when it isn't.

This article is based on the video RLHF from the Mark học AI channel. Watch the video (in Vietnamese) to see the animations.