ContentsThe library

Models That Carry a Memory

Add Instead of Multiply, and Decide How Much

Last timeWhy the Far Past Disappears

A gated update gives information a route through the chain that is addition rather than repeated transformation, and lets the model learn, per slot and per step, how much to keep.

The previous lesson left a specific problem. Information carried across many

steps is multiplied by a factor each time, the factor is a property of the

weights, and anything other than exactly one ruins it. The fix is to stop

multiplying.

FIG 1A gated update: keep some, write some
the keep gate: a number between zero and one for every slot, computed fresh at this step
the previous state, arriving unchanged rather than passed through a matrix
the write gate: how much of the proposed new content to admit, again per slot
the candidate: what this step would like to write, computed from the input and the previous state
Compare this with the plain update. There the previous state went through a matrix and a squashing function before becoming the new state. Here it is multiplied by a number the model chose and then added to. If that number is one, the slot passes through untouched, however many steps there are.

Add, do not transform

The important structural change is the addition. In the plain recurrence every

route from an early step to a late one goes through the same transformation

repeatedly, so the factor along it is some number raised to the number of steps.

Here there is a route that consists of multiplying by the keep gate and adding.

If the gate sits near one, the factor along that route is near one raised to the

number of steps, which is near one no matter how long the sequence is.

FIG 2Inside a gated cell
The previous state has two roles. It takes the straight route through the keep gate to the addition, and it also helps compute the gates and the candidate. The straight route is what carries long-range information; the other route is what makes the decisions.

The factor became something the model chooses

In a plain recurrence the per-step factor is fixed by the weights, so it applies

identically at every step of every sequence. A model cannot decide to hold on to

one particular thing, because it has no mechanism that distinguishes one step

from another.

A gate changes that. It is computed from the current state and input, so it is

different at every step and different for every slot. The model can learn to

drive one slot's keep gate to nearly one when something worth holding arrives,

leave it there while hundreds of irrelevant steps go past, and then drop it when

the thing is no longer needed.

The lesson stops here

3 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01Everything You Remember Has to Fit in the Same Box
  2. 02The Same Small Rule, Run a Thousand Timesopening only
  3. 03Almost Right, Four Hundred Times in a Rowopening only
  4. 04Add Instead of Multiply, and Decide How Muchyou are here
  5. 05A Thousand Waits That Did Not Have to Happenopening only
  6. 06Slots That Forget at Rates You Choseopening only
  7. 07A Rule That Reads What Arrived Before Decidingopening only
  8. 08The Bill That Grows and the Bill That Does Notopening only

Read alongside