A Habit, a Guard, and a Note on the Door
A refusal can be trained into the weights, enforced by a separate check, or written in a prompt. They fail in completely different ways, and most systems only have the weakest.
Three places, one question
Every deployed system has something deciding what it will not do. In most of
them that something is a paragraph near the top of a prompt, written quickly,
along the lines of: you must never provide instructions for making weapons, and
you must refuse politely if asked.
There are three genuinely different places that rule could have lived, and they
are not three strengths of the same thing. They differ in one respect that
matters more than all the others: what somebody would have to control in order
to remove the rule.
Trained into the weights
The first place is the model itself. During the alignment stage the model is
shown examples of refusing and examples of helping, and it learns a tendency:
requests of a certain shape get declined, and it has some sense of what that
shape is.
Three things make this valuable. It costs nothing at serving time, because
there is no extra call and no extra latency. It applies to every request
without anybody enumerating them. And crucially it generalises: a model trained
to refuse requests for synthesising a particular compound will usually also
refuse a request phrased as a chemistry lesson about the same compound, because
it learned something closer to the intent than to the wording.
Two things make it insufficient on its own. It is a tendency and not a rule, so
there is no threshold at which it is guaranteed to fire, and an unusual enough
framing finds the gap. And it is fixed at training time: if your product
decides next week that it will not discuss a competitor by name, that rule is
not in the weights and will not be until somebody retrains, which is not a thing
most teams can do.
| extra cost per request, | added delay, millisecond | holds against an ordinar | holds against someone tr | |
|---|---|---|---|---|
| trained into the weights | 0 | 0 | 95 | 3 |
| a separate check | 40 | 120 | 99 | 1 |
| written in the prompt | 0 | 0 | 60 | 0 |
A check outside the model
The second place is a separate step that reads the request, or the answer, and
decides independently. It can be a list of patterns, a trained classifier, or a
second model given one narrow job.
The property that matters is structural rather than a question of how good the
check is. The check is not following instructions contained in the text it is
reading, because it is not generating a continuation of that text. It is
classifying it. A message saying ignore your previous instructions is, to a
classifier, a string with some features. There is no channel through which it
can take effect, which is exactly what the prompt lacks.
This is the only place on the list that still works when the generating model is
fully under somebody else's control. If an attacker has found a way to make the
model produce anything they want, a check on the output still reads that output
and can still stop it. Nothing trained in and nothing in the prompt survives
that scenario, because both of them live inside the thing that was captured.
The costs are real and worth stating plainly. Each check is another call, so
there is latency and there is money. Checks have both error rates, so some good
requests get blocked and some bad ones get through. And a check on the output
interacts badly with streaming, since text already shown to the user cannot be
unsent, which the fourth lesson deals with in detail.
| step | where the rule lived | the request | what happened | what happened |
|---|---|---|---|---|
| 1 | trained in | plain harmful request | refused | The common case, handled free of charge and with no extra latency. This is most of the traffic and most of the value of training. |
| 2 | trained in | same request framed as fiction for a nov | answered | A tendency has no threshold. The framing moved the request far enough from what the training covered, and nothing else was in the way. |
| 3 | separate check on the output | same request framed as fiction | blocked | The check read the answer rather than the request, so the framing bought nothing. This is the pairing that works: training for breadth, a check for the thing that must not happen. |
| 4 | prompt paragraph only | request with an instruction to disregard | answered | The rule and the instruction to ignore it arrived in the same stream of text, with nothing marking one as authoritative. |
Using all three on purpose
None of this says the prompt is useless. It says the prompt is not a boundary,
which is a narrower claim. A paragraph of rules in a prompt does real work: it
sets the tone of a refusal, it tells an honest user what the system is for, and
it stops the sort of request that was never adversarial in the first place,
which is most of them. It just cannot be the thing standing between you and
somebody who wants past it.
So the design is not a choice between three options. It is an allocation.
Training handles breadth, which no list of rules can reach. A separate check
handles the specific things that must not happen, the ones where you would
rather lose a legitimate request than allow the bad one. The prompt handles tone
and expectation setting for people who are not attacking you.
The failure mode this course exists to prevent is having only the third, and
believing the paragraph is holding. It is not holding, and the next lesson shows
why that is a structural fact about how these systems read text rather than a
problem that better wording would solve.
Recap
- A rule can live in the weights, in a separate check outside the model, or in the prompt, and the three differ in what can remove them rather than in how strictly they are worded.
- Anything an attacker can edit cannot enforce a rule against that attacker, which disqualifies the prompt immediately and is the whole argument.
- Use all three deliberately: training for the common case, a separate check for the thing that must not happen, and the prompt for tone and for telling honest users what the system does.
This is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents