ContentsThe library

Many Specialists, One Model

Some Tokens Do Not Get In

Last timeWhen Everything Goes to One Place

Hardware wants fixed shapes, so each copy gets a fixed buffer. Tokens beyond it are silently skipped, and unused space is computed on anyway as though it were data.

Here is something that surprises people about trained sparse models: during

training, some tokens are simply not processed. Not by a worse copy, not by a

fallback, not at all. They pass through the layer and come out the other side

unchanged, and nothing anywhere reports it.

Why there is a buffer

The arithmetic in these models is done by hardware that is extremely good at one

thing: multiplying matrices whose shapes are known before the work begins and

whose contents sit in a contiguous block of memory. Everything about the

performance depends on that.

The routing decision is incompatible with it. How many tokens will choose the

third copy is not known until the chooser has run, it differs from batch to

batch, and it differs from copy to copy within one batch. So the arrangement that

suits the hardware is to declare in advance how much room each copy gets, build

a matrix of exactly that shape, and make reality fit.

FIG 1How big the room is
the number of tokens in the batch
how many copies each token is sent to, so that t times k is the number of token slots to be filled
the number of copies in the layer
the capacity factor, a number usually between one and two that decides how much slack each buffer has
the buffer size, the number of token slots each copy is given regardless of how many it actually receives
With a factor of one, every copy gets room for exactly its fair share, so any copy receiving more than average overflows. The factor is the only thing in this expression that anyone chooses, and it is the dial between two different kinds of waste.

Past the edge

Tokens are assigned to their chosen copies in some order, and when a copy's

buffer is full the next token that wanted it does not get in. What happens to

that token depends on nothing clever: it takes the residual path around the

layer and arrives at the next layer exactly as it entered this one.

The lesson stops here

4 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01Knowing More Without Doing More
  2. 02One Small Matrix Decides Everythingopening only
  3. 03Not Chemistry, Not Law, Not Medicineopening only
  4. 04The Favourite That Makes Itselfopening only
  5. 05Some Tokens Do Not Get Inyou are here
  6. 06The Exponential in the Middle of Everythingopening only
  7. 07The Model You Must Hold and the Model You Runopening only
  8. 08The Comparison That Actually Means Somethingopening only

Read alongside