Two seconds
You ask a bank's chatbot: my card is locked, how do I unlock it? Two seconds later, there's an answer. In those two seconds, your question goes through a whole assembly line.
STATION 1
Input guardrails
The first station is input guardrails, the gate on the way in. Here, the system hides your card number, and blocks off-topic or malicious questions.
STATION 2
The cloud
Then your question flies up to the cloud, to a data center where thousands of chips run the AI model. The answer is written there, not on your phone.
STATION 3
Output guardrails
The answer doesn't reach you right away. Output guardrails, the gate on the way out, check: does it promise something the policy doesn't allow, does it leak someone else's information? Real actions like unlocking a card wait for your confirmation.
STATION 4
Logging
Every step is logged, written into a record, like a plane's black box. If something goes wrong, engineers can trace what happened.
STATION 5
Monitoring
And monitoring tracks speed, cost and the rate of wrong answers on a dashboard. If anything looks unusual, an alert goes off.
EVALUATION
Change anything, test again
And every time the model changes or the prompt is edited, the system must run an evaluation on the old test set, and only goes live once it passes.
NEXT
Six stations ahead
In the next six parts, we stop at each station. Starting with where AI really lives: the cloud.
This article is based on the video Two seconds of a question from the Mark học AI channel. Watch the video (in Vietnamese) to see the animations.