ContentsThe library

Many Specialists, One Model

Not Chemistry, Not Law, Not Medicine

Last timeChoosing Which Parts Run

The copies do specialise, but almost never by subject. They divide up surface properties of single tokens, and the arrangements that work best are the ones that make that division easier to express.

The word chosen for these copies invites a particular picture: a panel of

specialists, one for medicine, one for law, one for code, and a receptionist who

knows which door to knock on. The picture is wrong in an interesting way, and

understanding why tells you what the arrangement is actually doing.

What a decision has to work with

Recall what the chooser sees. One token representation. Not a passage, not a

document, not a question. It cannot route by subject because at the moment of

the decision there is no subject in front of it, only a position in a sequence.

So if there is going to be division of labour, it has to be along properties a

single position can have. Those properties exist in quantity: whether this token

is punctuation, whether it is a digit, whether it is the start of a word or the

middle of one, whether it is a common function word or a rare fragment. That is

what measurement finds.

FIG 1Share of tokens of each kind reaching four copies, as a percentage
copy 1copy 2copy 3copy 4
punctuation41221918
digits12155419
common function words28242325
word fragments19441720
Three of the rows are clearly not uniform and one, the function words, is close to it. This is the shape of what gets found: real structure along surface categories, no structure along subject, and plenty of copies that are not specialised in any nameable way at all.

Why it does not happen by itself

Put eight copies in a layer, train, and measure. Typically the copies look

remarkably alike. There is a reason, and it is not a defect in the training.

Each copy receives a broad mixture of tokens, because the chooser starts out

near-random and keeps sending a wide spread to everybody. Each copy is then

trained to do as well as possible on that spread. The best thing a block can be,

when given a broad spread of inputs and asked to do well on all of them, is a

competent general-purpose block. Nothing anywhere in the objective says be

narrow. So the copies converge toward each other, and the model ends up holding

eight near-duplicates, which is a poor use of eight times the parameters.

The lesson stops here

4 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01Knowing More Without Doing More
  2. 02One Small Matrix Decides Everythingopening only
  3. 03Not Chemistry, Not Law, Not Medicineyou are here
  4. 04The Favourite That Makes Itselfopening only
  5. 05Some Tokens Do Not Get Inopening only
  6. 06The Exponential in the Middle of Everythingopening only
  7. 07The Model You Must Hold and the Model You Runopening only
  8. 08The Comparison That Actually Means Somethingopening only

Read alongside