ContentsThe library

What a Model Refuses, and Why

Build It So the Worst Case Is Boring

Last timeInstructions Hidden in the Data

Assume the model is fully under somebody else control and ask what it could do. The answer is a property of your tool definitions, not of the model, and you control it.

Assume it is captured

Every lesson so far has shown a way text reaches the model and changes what it

does. The prompt is one channel. Retrieved documents are a channel. Tool

results are a channel. The defences range from helpful to very helpful and none

of them is a guarantee.

So stop trying to guarantee it. Adopt the assumption instead: the model does

whatever an attacker wants. Now ask what happens.

This is not a counsel of despair, it is the normal way an untrusted component

is handled in any other part of a system. A browser cannot be trusted, so the

server validates. A client library cannot be trusted, so the API checks

permissions. Nobody finds this controversial, and the model is in exactly that

position: a component that processes attacker-influenced input and produces

requests your code then acts on.

The useful property of this assumption is that it makes the question

answerable. What could a captured model do is not a matter of opinion. It is

determined entirely by your tool definitions, which you wrote, and you can read

them.

The three powers that matter

Go through the tools and sort them by one question: if this is called with the

worst arguments an attacker could choose, is the result recoverable.

FIG 1Tools sorted by what the worst call does
model can call itdata or money leavescannot be undonerecoverable in minutes
search the knowledge bas1001
send an email1100
delete a record1110
issue a refund1110
grant a user a role1010
The first row is the shape every tool should aim for: callable, nothing leaves, nothing permanent. The four below it each have at least one of the three dangerous powers, and the marked cells are the ones that turn an incident into a disclosure, a loss or an unrecoverable change.

Sending data outward is the one most teams underestimate, because the tool

rarely looks like an exfiltration tool. An email sender is one. A webhook

caller is one. A tool that fetches a URL is one, because the URL can carry data

in it. A tool that writes to a shared document is one if anybody else can read

that document.

Writing irreversibly is the one that produces the worst Monday morning. Deletes

without a recycle bin, status changes that trigger downstream systems, messages

sent to customers.

Spending or granting is the one with a number attached. Refunds, purchases,

role assignments, quota increases.

The lesson stops here

4 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01A Habit, a Guard, and a Note on the Door
  2. 02Your Rules and Their Attack Arrive in the Same Envelopeopening only
  3. 03Two Mistakes, One Dial, and You Have to Chooseopening only
  4. 04The Only Layer That Sees What You Are About to Sayopening only
  5. 05The Attacker Does Not Have to Be the One Typingopening only
  6. 06Build It So the Worst Case Is Boringyou are here
  7. 07Nobody Writes In to Report a Refusalopening only
  8. 08If Two People Cannot Apply It, It Is Not a Ruleopening only

Read alongside