Mark học AI

Video #130 · AI on your own computer · Part 2/3

What is quantization? Shrinking an AI model to fit your laptop

You download an open model and the file is 16 GB, and your laptop runs out of memory? Bit explains quantization.

Watch videoVideo in Vietnamese

The model is too big

Mark tries downloading an open model, like in video 114, but the file is 16 GB and his laptop runs out of memory.

QUANTIZATION

Rounding the numbers

Bit explains: a model is billions of numbers. Quantization rounds those numbers so they take less space. Storing each number in 4 bits instead of 16 makes the model about four times lighter.

EXAMPLE

Like compressing a photo

It's like compressing a photo when you send it in a chat: the file gets much smaller, but it looks almost the same.

TRADE-OFF

A little less accuracy

In return, the model loses a bit of accuracy. 4 bits usually still works well, but squeeze it too hard and the answers get much worse.

LABELS

Reading the file names

When you download a model, you'll see labels like Q4 or Q8. That's the number of bits after compression: the smaller the number, the lighter the file.

This article is based on the video Quantization, shrinking a model to fit a laptop from the Mark học AI channel. Watch the video (in Vietnamese) to see the animations.