Not Chemistry, Not Law, Not Medicine
Last timeChoosing Which Parts Run
The copies do specialise, but almost never by subject. They divide up surface properties of single tokens, and the arrangements that work best are the ones that make that division easier to express.
The word chosen for these copies invites a particular picture: a panel of
specialists, one for medicine, one for law, one for code, and a receptionist who
knows which door to knock on. The picture is wrong in an interesting way, and
understanding why tells you what the arrangement is actually doing.
What a decision has to work with
Recall what the chooser sees. One token representation. Not a passage, not a
document, not a question. It cannot route by subject because at the moment of
the decision there is no subject in front of it, only a position in a sequence.
So if there is going to be division of labour, it has to be along properties a
single position can have. Those properties exist in quantity: whether this token
is punctuation, whether it is a digit, whether it is the start of a word or the
middle of one, whether it is a common function word or a rare fragment. That is
what measurement finds.
| copy 1 | copy 2 | copy 3 | copy 4 | |
|---|---|---|---|---|
| punctuation | 41 | 22 | 19 | 18 |
| digits | 12 | 15 | 54 | 19 |
| common function words | 28 | 24 | 23 | 25 |
| word fragments | 19 | 44 | 17 | 20 |
Why it does not happen by itself
Put eight copies in a layer, train, and measure. Typically the copies look
remarkably alike. There is a reason, and it is not a defect in the training.
Each copy receives a broad mixture of tokens, because the chooser starts out
near-random and keeps sending a wide spread to everybody. Each copy is then
trained to do as well as possible on that spread. The best thing a block can be,
when given a broad spread of inputs and asked to do well on all of them, is a
competent general-purpose block. Nothing anywhere in the objective says be
narrow. So the copies converge toward each other, and the model ends up holding
eight near-duplicates, which is a poor use of eight times the parameters.
The lesson stops here
4 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents