ContentsThe library

Keeping a Secret in Production

Both Values Have to Work for a While

Last timeDeciding Who May Read It

Rotation is a sequencing problem, not a security problem. Get the order right and nothing notices; get it wrong and everything fails at once, which is why nobody rotates.

Why both have to work

The reason most credentials in most systems have never been changed is not

laziness. It is that the obvious way to change one causes an outage, somebody

tried it once, and the organisation learned to leave it alone.

The obvious way is to replace the value. Update it in the store, and the next

thing that reads it gets the new one. The problem is the word next. Holders of

a credential do not all read it again at the same moment. A process that

started last Tuesday has the old value in memory and will keep it until it

restarts. A cached copy expires on its own schedule. A scheduled job will run

tomorrow morning with whatever it finds then. A partner system read it at

integration time and has it in a configuration file.

There is no instant at which everybody switches. So either both values work for

a period, or some holders are broken for that period.

FIG 1A rotation with an overlap, in days
2.557.51012.50new value created1both accepted4all services redeployed11old value unused for a week14old value removeddays from the start of the rotation
Fourteen days, and the long stretch between day four and day eleven is the part people want to cut. It is the part that catches the monthly job, the partner who deploys on Thursdays, and the instance that nobody knew was still running. Cutting it is how a rotation becomes an outage.

The order that works

Five steps, and the confirmations between them are the actual work.

Create the new value and register it alongside the old one, so the system

accepting the credential accepts either. Nothing has changed for any holder

yet, and nothing can break.

Deploy the new value to every holder. This takes as long as it takes, which

with scheduled jobs and partner systems is days rather than minutes.

Confirm that the new value is actually being used. Not that it was deployed:

that requests are arriving authenticated with it. This is the step that catches

the service which read the credential at startup three weeks ago and has not

restarted.

Confirm that the old value is no longer being used, by watching for any request

that still presents it. This is the step everybody skips, and it is the one

that turns the final step from a gamble into a formality.

Remove the old value, and confirm it now fails.

FIG 2The rotation sequence, with the checks that make it safe
Eleven nodes and two loops, and the loops are the point. The sequence is not a procedure that runs once; it is a procedure that waits at two gates until observation says it may continue. Teams that convert it into a linear checklist reintroduce the outage.
FIG 3The rotation runbook, written before you need it
plaintext
FOR THIS CREDENTIAL, WRITE DOWN NOW
  can the accepting system hold two values at once
  if not, what is the plan instead
  who holds it: services, jobs, partners, people
  how long until the slowest holder picks up a change
  how would you see it arriving in requests
  how would you see the old one still arriving

THE OVERLAP WINDOW IS
  the slowest holder, plus a full monthly cycle, plus a week

DURING AN EMERGENCY
  the same steps, with the window compressed deliberately
  and with somebody writing down what broke
  because that list is next quarters work
Written once per credential, this fits on half a page and turns a midnight rotation into a procedure somebody can follow while tired. The line about the window is the one worth arguing over, and the honest answer is always longer than the first estimate.

The ones that cannot overlap

Some credentials accept only one value at a time, and no amount of sequencing

changes that. These need a plan made in advance, because discovering the

limitation during an emergency is how a bad night becomes a long one.

A database account with one password is the common case. The account has a

password; setting a new one invalidates the old one at that instant. The

overlap has to be built elsewhere: create a second account, move holders to it,

then retire the first. That is a rotation of accounts rather than of passwords,

and it works, and it requires the privileges to have been granted to both.

A webhook configuration with a single secret field is the same problem in

somebody elses system. Some providers offer two slots precisely for this; where

they do not, the change is a brief window of rejected messages, and the plan is

to pick the quiet hour and to make sure the sender retries.

A shared symmetric key between two systems has to be changed on both sides at

once unless the protocol allows a key identifier, which is exactly what key

identifiers are for and why they are worth insisting on at design time.

The lesson stops here

2 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01Longer Than the List You Would Write From Memory
  2. 02Deleting It Does Not Undo Itopening only
  3. 03Two Questions Decide Every Place a Secret Can Sitopening only
  4. 04Follow the Chain Until It Stopsopening only
  5. 05One Component Falls, and Then Whatopening only
  6. 06Both Values Have to Work for a Whileyou are here
  7. 07Nobody Published It and It Is Publishedopening only
  8. 08The Order Is Not the One That Feels Urgentopening only

Read alongside