The Same Small Rule, Run a Thousand Times
Last timeA Fixed Summary Carried Forward
A recurrent model is one modest function applied repeatedly. Unrolling it shows both where its generality comes from and why a long sequence turns it into an extremely deep network.
The previous lesson said the state update is some function of the previous state
and the current input. This lesson makes that function concrete and then asks
what happens when you run it a thousand times.
- a squashing function that holds every number between minus one and one, which keeps the state from growing without limit
- the previous state passed through a matrix: this is where the model decides what to keep, mix or discard
- the current input passed through a different matrix, putting it into the same space as the state so the two can be added
- a bias, so the rule is not forced to send a zero state and zero input to zero
One step
The squashing function deserves attention. Without it the whole chain would be
one long multiplication by the same matrix, which either blows up or collapses
depending on that matrix. With it the state is kept inside a bounded region,
which prevents the blowing up and makes the collapsing worse. The next lesson is
about that trade.
It is worth noticing how little is in this line. There is no mechanism for
deciding that a particular input is important and should be protected. There is
no way for the rule to behave differently at step five than at step five
thousand. There is no place to store something aside for later. The rule gets
the state and the input, mixes them in a fixed way, and squashes the result.
Whatever selective memory the model appears to have must be an emergent property
of one matrix applied over and over, which is asking a great deal of it.
| step | step | a slot of the state | share of that slot owed to the first input | what happened |
|---|---|---|---|---|
| 1 | 1 | 0.62 | 100 | After the first step the state is entirely a function of the first input and the starting state. |
| 2 | 2 | 0.48 | 41 | The second input arrives and takes a large share of the slot. The first input is still clearly present. |
| 3 | 3 | 0.55 | 17 | A third arrival. Each update mixes the existing contents with something new, so earlier shares shrink multiplicatively rather than linearly. |
| 4 | 4 | 0.51 | 7 | Four steps in, the first input accounts for under a tenth of this slot, and the sequence has barely started. |
Depth measured in time
A stack of layers has a depth fixed when the model is built. A recurrence has a
depth equal to the length of its input, because an early input really does pass
through one application of the rule per subsequent position before it can
influence anything at the end.
The lesson stops here
2 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents