Round the Finished Model, or Train One That Expects It
Last timeWhere the Precision Still Matters
Rounding a trained model takes an hour and a few hundred sample inputs. Training one that knows it will be rounded costs a run and buys about two bits.
Everything so far has been applied to a model that was already finished. Take
the trained weights, choose step sizes, round, ship. That is one of two
families, and it is the cheap one. The other trains the model with the rounding
already in the picture, so the weights land somewhere that tolerates it.
The difference in cost is large and the difference in result is real. This
lesson is about how much of each.
The cheap method, and the set it depends on
Rounding a finished model needs no labels, no gradients and no training
infrastructure. For the weights there is nothing to decide beyond the step
sizes, since the weights are sitting there and can be inspected directly.
The activations are the complication. Their ranges are not visible in the
stored model, because activations do not exist until something is run through
it. So a few hundred inputs are pushed through, the values each part produces
are recorded, and the ranges are set from what was seen. Those inputs are the
calibration set, and they are the part of this method that quietly goes wrong.
| step | What the calibration inputs were | Inputs used | Quality on held-out traffic, points lost | What went wrong | what happened |
|---|---|---|---|---|---|
| 1 | Random web text | 512 | 4.1 | ranges set by text unlike the traffic | Convenient and wrong. The activation ranges a general corpus produces are not the ones a support assistant produces, and the clipping lands on the values that matter. |
| 2 | Real traffic, English only | 512 | 2.3 | other languages clip badly | Better, and it encoded a blind spot. Requests in other languages drive channels the set never visited, so their ranges are too narrow. |
| 3 | Real traffic, sampled across languages a | 512 | 0.6 | nothing identified | The set now looks like the traffic. This is the whole of the advice and it is routinely skipped in favour of whatever text was easy to obtain. |
| 4 | Real traffic, sampled across languages a | 32 | 1.9 | too few inputs to see the tails | The right distribution and not enough of it. The extreme channels appear rarely, so a small sample misses how far they reach. |
Two or three hundred inputs is usually enough once the distribution is right,
and a thousand is plenty. What matters far more than the count is that they
look like what the system will actually see: the same languages, the same task
mix, the same prompt lengths, including the awkward tail.
The lesson stops here
4 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents