ContentsThe library

Models That Carry a Memory

Almost Right, Four Hundred Times in a Row

Last timeOne Step, Applied Over and Over

Repeating one update rule means multiplying by roughly the same thing over and over, and the result either fades to nothing or grows past what can be represented. There is almost no middle.

The previous lesson ended with an uncomfortable observation: a recurrence

applies one rule hundreds of times in a row, and nothing corrects it between

applications. This lesson works out what that does.

Powers, not sums

Suppose each application of the rule leaves a fraction of whatever was carried

forward. Call that fraction the factor. After two steps the surviving amount has

been multiplied by the factor twice, after ten steps ten times, after four

hundred steps four hundred times. The quantity is a power, and powers do not

behave like the intuition people carry from addition.

FIG 1Percentage of an early input still present after a given number of steps
after 10 stepsafter 50after 200after 1000
a rule keeping 99 percen90.460.513.40.0
keeping 95 percent59.97.70.00.0
keeping 90 percent34.90.50.00.0
keeping 70 percent2.80.00.00.0
Read the second row. A rule that keeps ninety five percent of what it had sounds almost perfectly retentive, and the marked cell says that after two hundred steps nothing measurable survives. The top row, at ninety nine percent, is the only one that reaches two hundred steps at all, and it is gone by a thousand.
FIG 2The two things repeated multiplication can do
0.0015.0030.0045.0060.001.050.8100.5150.3200.0steps
keeping 95 percent each stepkeeping 102 percent each step
One curve is on the floor within a hundred steps and the other is heading for numbers too large to store. Both rules are within a few percent of neutral at a single step. There is no setting between them that stays flat for two hundred steps except exactly neutral, and nothing in ordinary training holds a rule exactly there.

It is worth being precise about what the factor is, because in a real model it

is not one number. The rule is a matrix, and what matters is how much it

stretches or shrinks the directions that carry the information you care about.

Some directions may be preserved while others fade. But the argument survives

the generalisation intact: whatever a direction's factor is, it gets raised to

the power of the number of steps, and the only factor that neither fades nor

grows is exactly one. A matrix that preserves every direction exactly is a very

particular kind of matrix, and gradient descent wandering through weight space

has no reason to find one and no reason to stay there once it has.

The lesson stops here

4 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01Everything You Remember Has to Fit in the Same Box
  2. 02The Same Small Rule, Run a Thousand Timesopening only
  3. 03Almost Right, Four Hundred Times in a Rowyou are here
  4. 04Add Instead of Multiply, and Decide How Muchopening only
  5. 05A Thousand Waits That Did Not Have to Happenopening only
  6. 06Slots That Forget at Rates You Choseopening only
  7. 07A Rule That Reads What Arrived Before Decidingopening only
  8. 08The Bill That Grows and the Bill That Does Notopening only

Read alongside