Mark học AI

Video #82 · Behind the scenes of an AI system · Part 4/7

What are output guardrails? How chatbots check answers before sending

AI has finished its answer, but it isn't sent yet: it has to pass output guardrails, the gate on the way out.

Follow on YouTubeComing soonVideo in Vietnamese

Written, but not sent yet

AI has finished writing its answer. But it isn't sent right away. It has to pass output guardrails, the gate on the way out.

CHECK 1

Is it making promises?

Check one: does it promise something the policy doesn't allow? As in video 61, an airline had to pay because its chatbot made up a refund policy. Guardrails compare the answer with the policy documents, and strike any line without a source.

CHECK 2

Leaky, rude, dangerous

Check two: does it leak someone else's information, or say anything rude or dangerous?

CHECK 3

Citations

Check three: does the answer have citations pointing to the right documents? With sources, customers can check for themselves.

HUMAN IN THE LOOP

A person clicks confirm

For real actions, like unlocking a card or sending money, AI only gets things ready. The user has to click confirm, and sometimes a staff member has to approve too. This is called human in the loop.

TWO GATES

Like airport security

Together, the two gates are like airport security: scan on the way in, check on the way out. If one gate misses something, the other can still catch it.

NEXT

Will a new model break it?

The guardrails are in place. But will switching to a new model break things? Part five: evaluation.

This article is based on the video Output guardrails: the gate on the way out from the Mark học AI channel. Watch the video (in Vietnamese) to see the animations.