Keeping a Secret in Production
Deleting It Does Not Undo It
Last timeWhat Counts as a Secret
Everybody knows not to commit a secret. Fewer people know why removing it afterwards changes nothing, and that knowing why decides what you do in the first ten minutes.
What a history is for
A version control system has one defining property: it keeps every state a file
has ever been in, so that any of them can be recovered later. That is not a
side effect of the implementation. It is the entire reason the tool exists, and
every design decision in it serves that property.
Now consider deleting a committed password. You edit the file, remove the line,
and commit. What you have done is add a new state in which the line is absent.
The earlier state, in which the line is present, is untouched, still stored,
still addressable by its identifier, and still served to anybody who asks for
it. You have not removed anything. You have appended.
This is worth sitting with, because the intuition from every other kind of file
points the other way. Deleting a line from a document removes the line. Here it
does not, and the tool is working correctly.
Who already has a copy
The value exists in more places than the repository, and most of them are
outside your control within the hour.
Every clone is a full copy of the history, which is what makes the tool fast
and what makes this problem permanent. A teammate who pulled on Tuesday has the
value on their laptop whether or not they ever look at it. A fork has it.
A mirror has it. A caching proxy in front of the hosting service may have it.
The build system has it twice: once in the checkout, and once in whatever the
build printed, which is a second publication under different retention rules
and usually broader access. Logs are frequently readable by people who cannot
read the repository at all.
And then there are the collectors. Watching a public hosting service for newly
published credentials is inexpensive, fully automated, and continuous, and the
collected values are tested against common services immediately. The relevant
unit of time is minutes.
Scanning finds it second
Scanning your own repositories for credentials is worth doing. It catches
mistakes, it catches them repeatedly, and it is cheap. It is not a defence, and
treating it as one produces exactly the wrong response at the wrong speed.
The asymmetry is structural. The collectors are continuous, automatic, and
watching everything published anywhere. Your scanner runs on a schedule, on
your own repositories, and reports to a channel somebody reads eventually. You
are not racing them. You are finding out second, by design, and no amount of
running the scan more often changes which side of the race you are on.
The useful place for a scan is before the push rather than after it. A check
that refuses the commit prevents the publication entirely, and that is a
different kind of tool from one that reports a publication that has happened.
The lesson stops here
3 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents