Adapting a Model You Did Not Train
Spending a Fixed Budget in the Right Places
Last timeThe Change Is Smaller Than the Model
Given a number of trainable parameters to spend, the decisions that matter are which weights get a correction, how narrow each one is, and how loudly it speaks.
The previous lesson established the construction and one dial. In practice there
are three decisions, and the first of them is the one teams get wrong most
often: which weights receive a correction at all.
Where they attach
A transformer block has several weight matrices. Four of them belong to
attention, and two or three larger ones belong to the feed-forward part, which
holds most of the parameters.
| 0.1M trained | 0.3M trained | 1M trained | 3M trained | |
|---|---|---|---|---|
| query and value only | 0.82 | 0.85 | 0.86 | 0.86 |
| all four attention proje | 0.84 | 0.88 | 0.89 | 0.89 |
| attention and feed-forwa | 0.86 | 0.90 | 0.91 | 0.91 |
Breadth beats width
That observation generalises into the single most useful rule here. Given a
fixed number of trainable parameters, spread them over as many weight matrices
as you can and make each bottleneck as narrow as that requires.
The lesson stops here
6 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents