The Only Layer That Sees What You Are About to Say
Last timeChecking What Comes In
A check on the answer catches what no check on the request could predict, and it collides head-on with streaming, which is why most teams skip it.
The failures that live in the answer
The previous lesson ended on a limit. An input check reads the request, so it
can only catch what is wrong with the request. A large category of problem is
not wrong with the request at all.
Consider four cases, all of which begin with a question any reasonable check
would allow.
| step | the request | why the input check allowed it | what came out | what happened |
|---|---|---|---|---|
| 1 | what is our refund window for enterprise | an ordinary support question | the answer, plus an internal note about | The retrieval layer pulled an internal document into the context and the model used it. The request was fine and the failure is entirely in the answer. |
| 2 | can you summarise my last conversation w | a legitimate request from an authenticat | a summary of a different user conversati | A wrong identifier somewhere upstream. No input check could have known, because the request was correct and the system answered a different question than it was asked. |
| 3 | what does section 4 of our contract say | a routine document question | a confident paraphrase of a section that | Invented, fluently, with no hedging. Visible only by checking the answer against the source, which is a check and not a prompt instruction. |
| 4 | write a polite note declining a vendor | entirely benign | a note quoting the internal budget figur | The context contained something the user was allowed to see and the recipient was not. Correct for the user, wrong for the destination. |
The common structure: the request was fine, the context was not, or the
plumbing was not. The only place the problem exists is in the text about to be
sent, so the only place to catch it is there.
The collision with streaming
Here is why most teams do not do this.
Streaming shows the answer as it is generated, a few words at a time, because
a visible first word in two hundred milliseconds feels like a different product
from a blank screen for six seconds. Every serious chat interface streams.
A check needs text to judge. Judging the first four words of an answer is
mostly guesswork. And text already shown to the user cannot be unsent: you can
cover it, replace it, apologise for it, but they saw it.
The lesson stops here
3 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents