The model is too big
Mark tries downloading an open model, like in video 114, but the file is 16 GB and his laptop runs out of memory.
QUANTIZATION
Rounding the numbers
Bit explains: a model is billions of numbers. Quantization rounds those numbers so they take less space. Storing each number in 4 bits instead of 16 makes the model about four times lighter.
EXAMPLE
Like compressing a photo
It's like compressing a photo when you send it in a chat: the file gets much smaller, but it looks almost the same.
TRADE-OFF
A little less accuracy
In return, the model loses a bit of accuracy. 4 bits usually still works well, but squeeze it too hard and the answers get much worse.
LABELS
Reading the file names
When you download a model, you'll see labels like Q4 or Q8. That's the number of bits after compression: the smaller the number, the lighter the file.
This article is based on the video Quantization, shrinking a model to fit a laptop from the Mark học AI channel. Watch the video (in Vietnamese) to see the animations.