Big but still fast?
Mark reads the specs of a Mixture of Experts model: a huge number of parameters, yet it still runs fast. He wants to know why.
MoE
A team of experts
Bit explains: Mixture of Experts, MoE for short, splits a model into many small experts, plus a coordinator called the router.
HOW IT RUNS
Only a few are called each time
For each token that comes in, the router calls only the few best-fitting experts, and the others rest.
THE BENEFIT
A big store, a light load
That way, the model has a huge store of knowledge, but only runs a small part each time, so it stays fast and cheap.
GETTING IT RIGHT
Not split by school subject
Note: the experts aren't neatly split into math, literature or history. They divide the work themselves while learning, in ways people find hard to name.
REMEMBER
When you see this word
When you see this word, think: a big model that only uses a few experts at a time, so it's fast. Many powerful models today work this way.
This article is based on the video What is Mixture of Experts (MoE)? from the Mark học AI channel. Watch the video (in Vietnamese) to see the animations.