Adapting a Model You Did Not Train
The Abilities Nobody Was Trying to Change
Last timeChanging Every Weight
Adapting a model on narrow data moves weights that other abilities depended on. The loss is real, it is invisible to the metric being optimised, and it has to be measured on purpose.
Adaptation is usually described by what it adds. The more useful description is
what it removes, because that part happens without anyone asking for it and
without appearing in any number the team is watching.
Why it happens at all
The mechanism is not subtle and it follows directly from how training works. A
weight in a language model is not dedicated to one ability. The same weight
participates in understanding negation, in handling a second language, in
declining a harmful request, and in whatever your task needs.
What goes first
The pattern is predictable from the adaptation data. Whatever is absent from it
is what degrades, because nothing in the training signal is pushing back.
The lesson stops here
7 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents