The Mistakes Behind Most Breaches
One Mistake, Wearing Different Clothes
Database injection, command injection, scripting in a page and template injection are not four problems. They are one problem at four boundaries, and knowing that gives you one fix.
The single cause
There is a list of injection vulnerabilities, and lists are how this topic
is usually taught: one for databases, one for shell commands, one for
markup in a page, one for template engines, one for directory services,
one for the query language your search index uses.
Treat it as a list and you will learn six sets of rules, remember four of
them, and be defenceless against the seventh, which arrives next year
attached to a technology nobody has heard of yet.
There is one cause. A string is assembled out of a part you wrote and a
part that came from outside, and then it is handed to something that
parses the whole string looking for structure. The parser has no way to
know which characters came from where, because the string does not record
it. By the time the quotation mark arrives, it is just a quotation mark.
| string assembled from tw | handed to a parser, 0 or | parser cannot tell the p | fixed by sending the val | |
|---|---|---|---|---|
| a database query built b | 1 | 1 | 1 | 1 |
| a shell command built fr | 1 | 1 | 1 | 1 |
| markup built from a comm | 1 | 1 | 1 | 1 |
| a template rendered from | 1 | 1 | 1 | 1 |
| a directory lookup built | 1 | 1 | 1 | 1 |
| a search query built fro | 1 | 1 | 1 | 1 |
Finding the boundary
So the useful skill is not knowing six lists. It is being able to find,
in your own code, every place where a string you partly built is given to
something that parses it.
They are easy to recognise once you are looking. The function being called
takes a whole command, a whole query, a whole document, rather than a
structure with slots. Somewhere just above it, a string was built.
interpreter | the shape of the mistake | the shape of the fix
--------------|-------------------------------|------------------------------
database | query text built by joining | a statement with placeholders
shell | a command line built by joining | a program plus an argument list
a page | markup built by joining | set the text of a node
a template | a template built from input | a fixed template, data passed in
a lookup | a filter built by joining | a filter with bound values
a search | a query built by joining | the clients structured query
how to find them:
search for the join operator next to a call that
takes one whole string, in any of the rows above
what not to search for:
lists of characters, which tell you nothing
about which direction the data was goingWhy escaping is not the fix
The usual response is to clean the input: strip or escape the characters
that mean something to the interpreter. It is the wrong tool and it fails
in a predictable way.
Escaping is correct only relative to a context. A value escaped for a
database is not escaped for markup. A value escaped for markup is wrong
inside a script block within that markup, which is a different context
nested inside the first, with different rules. A value escaped for a
shell is wrong in the second shell that the first one invokes.
| step | where it goes | escaped for | safe there | why or why not | what happened |
|---|---|---|---|---|---|
| 1 | into the database | the database | yes | the value was bound rather than joined, | The first boundary is usually the one people fix, because it has the famous name. It is also the one a library makes easy. |
| 2 | out into a page | the database | no | database escaping means nothing to a mar | The value was cleaned for the wrong consumer. It travelled through storage, which erased the context it was cleaned for, and nobody at the second boundary knew. |
| 3 | into an attribute in that page | markup | no | attribute context has different rules th | Nested contexts are where escaping loses even to careful people. The same value needs different treatment depending on where in the document it lands. |
| 4 | into a report sent to a shell tool | markup | no | the shell parses characters markup does | A path nobody designed. A reporting job reads the field and passes it to a command line, two years after the field was added. |
| 5 | bound as a value at each boundary | nothing | yes | no parser ever received it as part of a | The fixed version of all four rows. Note the second column: nothing was escaped anywhere, and the value may contain any characters at all. |
There is a second reason, which matters more in practice. Escaping puts
the burden on the person writing the line, every time, forever. The fix
has to be reapplied by the next person who edits that string, including
the one doing something unrelated at two in the morning. Any defence that
requires correct behaviour at every call site, indefinitely, is a defence
that will fail.
- sites where an injection could occur
- distinct interpreters in the system: databases, shells, markup, templates
- call sites per interpreter, which grows with every feature
Handing the value across separately
State the fix in a way that covers interpreters that do not exist yet.
Send the structure and the values on separate channels, so that the
interpreter receives your structure as structure and the value as a
value, and no parsing step ever has the opportunity to confuse them.
That is one sentence and it instantiates everywhere. A prepared statement
with placeholders, where the query text is fixed and the values are bound
afterwards. A program invoked with an argument list rather than a command
line, so no shell parses anything. Setting the text content of an element
rather than assembling markup. A template compiled from a fixed source,
with data supplied to it.
Two cases do not fit and are worth naming, because pretending otherwise is
how people get hurt. Some things genuinely cannot be bound: a table name,
a column to sort by, a flag passed to a command. For those the answer is
a permitted list of known-good values, mapping a supplied token to one of
a fixed set, never passing the supplied text through. And some systems
must accept rich text from users, where markup is the point. That is not
solved by escaping either; it needs a parser that builds a document and
emits only known-safe elements, which is a library you adopt rather than
a function you write.
Validation still has a place and it is not this one. Check that a figure
is a number and that a date is a date, because that is correctness and it
catches mistakes. Just do not let it stand in for the structural fix: a
value that passed every validation rule is still a value, and if it is
concatenated into a string that something parses, the problem is
unchanged.
The remaining lessons are the other recurring mistakes, and two of them
are this one again at different boundaries: content that gets interpreted
in a browser, and an address that gets fetched by your own server.
Recap
- Every injection is the same event: a string is assembled from a trusted part and an untrusted part, and then handed to something that parses the whole string for structure.
- Escaping is a patch on a design that mixes code and data. It works until a context changes, and the context changes whenever somebody edits the surrounding string.
- The structural fix is to hand the untrusted part across separately, as a value, so that no parser ever sees it as something that could be structure.
This is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents