ContentsThe library

What a Model Refuses, and Why

A Habit, a Guard, and a Note on the Door

A refusal can be trained into the weights, enforced by a separate check, or written in a prompt. They fail in completely different ways, and most systems only have the weakest.

Three places, one question

Every deployed system has something deciding what it will not do. In most of

them that something is a paragraph near the top of a prompt, written quickly,

along the lines of: you must never provide instructions for making weapons, and

you must refuse politely if asked.

There are three genuinely different places that rule could have lived, and they

are not three strengths of the same thing. They differ in one respect that

matters more than all the others: what somebody would have to control in order

to remove the rule.

FIG 1Where a request meets each of the three
The two boxes outside the model are the ones an attacker cannot write into. Everything between them is text, and text is the attacker's medium. Which box a rule sits in decides what it survives.

Trained into the weights

The first place is the model itself. During the alignment stage the model is

shown examples of refusing and examples of helping, and it learns a tendency:

requests of a certain shape get declined, and it has some sense of what that

shape is.

Three things make this valuable. It costs nothing at serving time, because

there is no extra call and no extra latency. It applies to every request

without anybody enumerating them. And crucially it generalises: a model trained

to refuse requests for synthesising a particular compound will usually also

refuse a request phrased as a chemistry lesson about the same compound, because

it learned something closer to the intent than to the wording.

Two things make it insufficient on its own. It is a tendency and not a rule, so

there is no threshold at which it is guaranteed to fire, and an unusual enough

framing finds the gap. And it is fixed at training time: if your product

decides next week that it will not discuss a competitor by name, that rule is

not in the weights and will not be until somebody retrains, which is not a thing

most teams can do.

FIG 2The three places on four questions
extra cost per request, added delay, millisecondholds against an ordinarholds against someone tr
trained into the weights00953
a separate check40120991
written in the prompt00600
The last column is the one that matters and the first two are what people actually decide on. A prompt rule is free and instant and scores zero against anybody who is trying, which is a bad trade dressed up as an efficient one.

A check outside the model

The second place is a separate step that reads the request, or the answer, and

decides independently. It can be a list of patterns, a trained classifier, or a

second model given one narrow job.

The property that matters is structural rather than a question of how good the

check is. The check is not following instructions contained in the text it is

reading, because it is not generating a continuation of that text. It is

classifying it. A message saying ignore your previous instructions is, to a

classifier, a string with some features. There is no channel through which it

can take effect, which is exactly what the prompt lacks.

This is the only place on the list that still works when the generating model is

fully under somebody else's control. If an attacker has found a way to make the

model produce anything they want, a check on the output still reads that output

and can still stop it. Nothing trained in and nothing in the prompt survives

that scenario, because both of them live inside the thing that was captured.

The costs are real and worth stating plainly. Each check is another call, so

there is latency and there is money. Checks have both error rates, so some good

requests get blocked and some bad ones get through. And a check on the output

interacts badly with streaming, since text already shown to the user cannot be

unsent, which the fourth lesson deals with in detail.

FIG 3One request against each of the three
stepwhere the rule livedthe requestwhat happenedwhat happened
1trained inplain harmful requestrefusedThe common case, handled free of charge and with no extra latency. This is most of the traffic and most of the value of training.
2trained insame request framed as fiction for a novansweredA tendency has no threshold. The framing moved the request far enough from what the training covered, and nothing else was in the way.
3separate check on the outputsame request framed as fictionblockedThe check read the answer rather than the request, so the framing bought nothing. This is the pairing that works: training for breadth, a check for the thing that must not happen.
4prompt paragraph onlyrequest with an instruction to disregardansweredThe rule and the instruction to ignore it arrived in the same stream of text, with nothing marking one as authoritative.
4 steps
Four rows and the whole argument of the course. Training catches the ordinary case cheaply, the separate check catches what training missed, and the prompt catches a user who was not trying to get past it.

Using all three on purpose

None of this says the prompt is useless. It says the prompt is not a boundary,

which is a narrower claim. A paragraph of rules in a prompt does real work: it

sets the tone of a refusal, it tells an honest user what the system is for, and

it stops the sort of request that was never adversarial in the first place,

which is most of them. It just cannot be the thing standing between you and

somebody who wants past it.

FIG 4Where refusals actually came from, over a month of production traffic
Training does the bulk of the work by volume, which is why it feels sufficient. The checks account for a fifth of refusals and almost all of the ones that mattered, since they are what fired on the requests that got past the training.

So the design is not a choice between three options. It is an allocation.

Training handles breadth, which no list of rules can reach. A separate check

handles the specific things that must not happen, the ones where you would

rather lose a legitimate request than allow the bad one. The prompt handles tone

and expectation setting for people who are not attacking you.

The failure mode this course exists to prevent is having only the third, and

believing the paragraph is holding. It is not holding, and the next lesson shows

why that is a structural fact about how these systems read text rather than a

problem that better wording would solve.

Recap

  • A rule can live in the weights, in a separate check outside the model, or in the prompt, and the three differ in what can remove them rather than in how strictly they are worded.
  • Anything an attacker can edit cannot enforce a rule against that attacker, which disqualifies the prompt immediately and is the whole argument.
  • Use all three deliberately: training for the common case, a separate check for the thing that must not happen, and the prompt for tone and for telling honest users what the system does.

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

NextWhy a Prompt Is Not a Boundary →

The rest of this course

  1. 01A Habit, a Guard, and a Note on the Dooryou are here
  2. 02Your Rules and Their Attack Arrive in the Same Envelopeopening only
  3. 03Two Mistakes, One Dial, and You Have to Chooseopening only
  4. 04The Only Layer That Sees What You Are About to Sayopening only
  5. 05The Attacker Does Not Have to Be the One Typingopening only
  6. 06Build It So the Worst Case Is Boringopening only
  7. 07Nobody Writes In to Report a Refusalopening only
  8. 08If Two People Cannot Apply It, It Is Not a Ruleopening only

Read alongside