Mark học AI

Video #43 · How AI learns · Part 6/7

Why do bigger AI models get smarter? Scaling laws explained

Why has AI suddenly gotten so much better in the last few years? The secret is three sliders.

Watch videoVideo in Vietnamese

AI suddenly got much better

Last time: a machine has to learn to understand, not memorize. So why has AI suddenly gotten so much better in the last few years? The secret: bigger.

THREE SLIDERS

Three things get pushed up

Picture three sliders. Slider one: the number of knobs in the network. Slider two: the amount of data to learn from. Slider three: computing power, meaning the number of graphics chips, called GPUs.

SCALING LAWS

Push them up, the error goes down

People noticed that pushing all three up together makes the error drop quite steadily. These are called scaling laws. Thanks to them, people can roughly predict how much better a bigger model will be.

BALANCE

Keep them balanced

But they have to stay balanced. Lots of knobs with little data is like a big brain with just one book to read. Lots of data with few knobs is like a huge library with a tiny memory.

THE RACE

Racing to build data centers

That's why big companies race to build data centers with tens of thousands of chips, using a lot of electricity. It's also why big models are usually slower and pricier than small ones.

LIMITS

Bigger isn't magic

Getting one step better costs many times more, and the error only drops a little. So besides making models bigger, people also look for smarter ways to teach them.

PART 7

Does anyone understand the inside?

With something that big, does anyone really understand what's going on inside? See you in part 7, the last part.

This article is based on the video Why bigger means smarter from the Mark học AI channel. Watch the video (in Vietnamese) to see the animations.