2 seconds become 10 minutes
You type a question, press Enter, and two seconds later AI answers. This is our two-hundredth special episode, and we'll slow those two seconds down into ten minutes to see what happens inside. Mark's question is very ordinary: is it cold in Hanoi tomorrow, and what should he pack? The moment his finger hits Enter, the clock starts.
STAGE 1
The question doesn't travel alone
Stage one, the question doesn't travel alone: before sending, the app packs in the earlier messages of the chat, plus a hidden instruction called the system prompt, written by the app makers, which you never see. All of it together is the context. Remember the desk in video 9: AI only knows what's on the desk, and can't see anything off it.
STAGE 2
Up to the cloud: API, data center
Stage two, up to the cloud: the envelope flies across the internet, through a door called an API, to a data center, buildings full of servers with GPUs, chips built for AI math. At the entrance there are input guardrails that check the question for anything dangerous or sensitive. Mark's weather question passes easily.
STAGE 3
Cutting it up: tokens
Stage three, cutting it up: AI doesn't read words the way we do, so the question is cut into small pieces called tokens. Some are whole words, some are half a word or a punctuation mark. Vietnamese, with its accent marks, is usually cut into more pieces than English, so the same idea costs more tokens. Each piece gets a number, like a page number in a very thick dictionary.
STAGE 4
The map of meaning: embeddings
Stage four, placing them on a map of meaning: each token is turned into a long list of numbers called an embedding. It works like coordinates on a giant map where words with similar meanings sit close together. On that map, cold sits next to jacket and scarf, and Hanoi sits near Hoan Kiem Lake, the Old Quarter and autumn. That's how AI grasps meaning, not just the letters.
STAGE 5
The pieces look at each other: attention
Stage five, the pieces look at each other: this is the most important part, called attention. Each token looks at all the other tokens to find which ones matter most to it. The word for pack in the question pays strong attention to Hanoi and cold, so AI understands Mark needs clothes for a cold trip, not goods to sell. This looking happens again across dozens of layers, each one understanding a bit more. Billions of multiplications run in parallel on GPUs, all in a fraction of a second.
STAGE 6
A thinking draft
Stage six, a thinking draft: many models today are reasoning models that write a draft before answering. The draft is very honest: the question needs tomorrow's weather, but the model's knowledge stops when its training ended, so it doesn't know.
STAGE 7
Picking up a tool: tool calling
Stage seven, picking up a tool: instead of guessing, AI writes a special request: call the weather tool, place Hanoi, tomorrow. This is called tool calling. Interestingly, AI doesn't run the tool itself: an outside system reads the request, calls the weather service and puts the result on the desk: a cold front tomorrow, 16 to 24 degrees, maybe light rain in the afternoon. Now there's real data on the desk, which is also why AI with search answers questions about today far better than AI without it.
STAGE 8
Guessing the first word
Stage eight, guessing the first word: only now does AI start answering, and it answers by guessing. It scores tens of thousands of tokens that could start the reply. "Yes" gets forty percent, "Tomorrow" thirty, "According" fifteen. Then it picks one by those odds, like spinning a prize wheel with big and small slices. This randomness is set by the temperature. Because of this draw, asking the same question twice can give two different answers.
STAGE 9
Repeating word by word
Stage nine is repetition until the answer is done. The wheel stops on "Yes", that word is added to the end, and the whole sequence, now one piece longer, runs through all the attention layers again to guess the next piece. Mark's answer is about eighty tokens long, so the whole line runs about eighty times. Each word is sent to the phone as soon as it's ready, called streaming, which is why the answer appears word by word as if someone were typing.
STAGE 10
The exit, and where it can go wrong
Stage ten, the exit: before reaching you, the answer passes output guardrails that block harmful content or leaked information. But guardrails can't check whether it's true. Without the weather tool, AI would still guess a sentence that sounds very sure, just not based on real data. That's hallucination, and it comes from the very same word-guessing step.
2.00 SECONDS
Back in your hands
The clock stops at two seconds, and the screen shows the answer: a cold front reaches Hanoi tomorrow, it'll be chilly with maybe light rain in the afternoon, so bring a jacket, a scarf and a folding umbrella. Those two seconds held a context envelope, a trip to a data center, a few dozen tokens, dozens of attention layers, one tool call and about eighty word guesses. Every one of those costs electricity and money, which is why free apps usually limit how many messages you can send.
200 videos in one press of Enter
The channel's two hundred videos, wrapped up, are this one press of Enter: from tokens and the map of meaning to attention, tool calling and hallucination. Thank you for learning with Mark and Bit for two hundred episodes. Every stage today has its own video.
This article is based on the video Special episode #200: Dissecting one press of Enter from the Mark học AI channel. Watch the video (in Vietnamese) to see the animations.