AI suddenly got much better
Last time: a machine has to learn to understand, not memorize. So why has AI suddenly gotten so much better in the last few years? The secret: bigger.
THREE SLIDERS
Three things get pushed up
Picture three sliders. Slider one: the number of knobs in the network. Slider two: the amount of data to learn from. Slider three: computing power, meaning the number of graphics chips, called GPUs.
SCALING LAWS
Push them up, the error goes down
People noticed that pushing all three up together makes the error drop quite steadily. These are called scaling laws. Thanks to them, people can roughly predict how much better a bigger model will be.
BALANCE
Keep them balanced
But they have to stay balanced. Lots of knobs with little data is like a big brain with just one book to read. Lots of data with few knobs is like a huge library with a tiny memory.
THE RACE
Racing to build data centers
That's why big companies race to build data centers with tens of thousands of chips, using a lot of electricity. It's also why big models are usually slower and pricier than small ones.
LIMITS
Bigger isn't magic
Getting one step better costs many times more, and the error only drops a little. So besides making models bigger, people also look for smarter ways to teach them.
PART 7
Does anyone understand the inside?
With something that big, does anyone really understand what's going on inside? See you in part 7, the last part.
This article is based on the video Why bigger means smarter from the Mark học AI channel. Watch the video (in Vietnamese) to see the animations.