What's it thinking inside?
Last time: the bigger AI gets, the better it gets. But billions of knobs turn themselves. Does anyone understand what's going on inside?
BLACK BOX
A black box
People call AI a black box: a question goes in, an answer comes out, and in between there are just billions of numbers. The field of interpretability is shining a light into that box.
CONCEPTS
Finding concepts
Researchers found groups of neurons that light up together for one concept. For example, one group lights up whenever the Golden Gate Bridge comes up, in many languages or in photos.
TURNING IT UP
Turning up a concept
They can even turn it up. Turn this group up hard, and whatever AI says steers back to the bridge, and it even claims to be the bridge. So inside there are concepts we can read, not just meaningless numbers.
STILL FAR TO GO
Only one corner is lit so far
But the flashlight only lights a small corner. Most of the inside is still not understood. Understanding it is one way to make AI safer and more trustworthy.
RECAP
Seven parts in one picture
Looking back at the series: knobs, measuring the error, walking downhill, tracing mistakes back, understanding instead of memorizing, scaling up to get better, and shining a light into the black box.
NEW SERIES
How AI gets graded
That's how AI learns. But how do we know how good it is? See you in the new series: how AI gets graded. Just joining? Start from part 1.
This article is based on the video Looking into the black box from the Mark học AI channel. Watch the video (in Vietnamese) to see the animations.