Written, but not sent yet
AI has finished writing its answer. But it isn't sent right away. It has to pass output guardrails, the gate on the way out.
CHECK 1
Is it making promises?
Check one: does it promise something the policy doesn't allow? As in video 61, an airline had to pay because its chatbot made up a refund policy. Guardrails compare the answer with the policy documents, and strike any line without a source.
CHECK 2
Leaky, rude, dangerous
Check two: does it leak someone else's information, or say anything rude or dangerous?
CHECK 3
Citations
Check three: does the answer have citations pointing to the right documents? With sources, customers can check for themselves.
HUMAN IN THE LOOP
A person clicks confirm
For real actions, like unlocking a card or sending money, AI only gets things ready. The user has to click confirm, and sometimes a staff member has to approve too. This is called human in the loop.
TWO GATES
Like airport security
Together, the two gates are like airport security: scan on the way in, check on the way out. If one gate misses something, the other can still catch it.
NEXT
Will a new model break it?
The guardrails are in place. But will switching to a new model break things? Part five: evaluation.
This article is based on the video Output guardrails: the gate on the way out from the Mark học AI channel. Watch the video (in Vietnamese) to see the animations.