Mark học AI

Video #115 · 60-second AI glossary · Part 3/5

What is distillation in AI? A big model teaches a small one

Why are some tiny models, small enough to run on a phone, still pretty smart? One secret is distillation.

Follow on YouTubeComing soonVideo in Vietnamese

Small but smart?

Mark wonders why some tiny models, small enough to run on a phone, are still pretty smart.

DISTILLATION

A big teacher, a small student

One secret is distillation. A big model is the teacher, answering lots and lots of questions. A small model is the student, learning from how the teacher answers.

WHY

Learning the hesitation too

The student doesn't just learn the right answer. It also learns how much the teacher hesitated between answers. That information helps the student learn much faster.

THE RESULT

Small, fast, cheap

The result is a small, fast, cheap model that keeps most of the teacher's skill, especially at the tasks it was taught.

REMEMBER

When you see this word

When you see this word, think: a small model learning from a big one. That's why there are fast, cheap mini versions, as in video 18.

This article is based on the video Distillation: a big model teaches a small one from the Mark học AI channel. Watch the video (in Vietnamese) to see the animations.