The library

How models work

Be able to explain what changing a pretrained model actually involves, why a very small change is usually enough, what each cheap method alters, what adaptation destroys, and how to choose between adapting a model and the alternatives

Adapting a Model You Did Not Train

Almost nobody trains a model from nothing. The useful skill is changing one that already exists, cheaply enough to afford and carefully enough not to ruin it.

8 lessons, written and corrected before you arrived. Reading them here needs no account. The first reads the whole way through; the others open and then stop, because a page nobody owns cannot tell who is reading it. Starting the course gives you your own copy, where every idea has problems standing under it and you can ask about any sentence.

Start reading

  1. 01Almost All of the Work Is Already DoneA pretrained model already knows how language works and what things are called. What it lacks for your task is a small correction, and the whole subject is about paying only for that.
  2. 02The Obvious Approach and What It Costsopening onlyContinuing to train a model on your own data is the direct way to adapt it. It works, it sets the quality ceiling, and it needs several times the model size in memory to do.
  3. 03The Abilities Nobody Was Trying to Changeopening onlyAdapting a model on narrow data moves weights that other abilities depended on. The loss is real, it is invisible to the metric being optimised, and it has to be measured on purpose.
  4. 04Writing the Correction as Two Thin Sheetsopening onlyThe correction adaptation needs is far simpler than the model it corrects, so it can be written as a product of two thin matrices and trained with a thousandth of the parameters.
  5. 05Spending a Fixed Budget in the Right Placesopening onlyGiven a number of trainable parameters to spend, the decisions that matter are which weights get a correction, how narrow each one is, and how loudly it speaks.
  6. 06Four Other Places to Put the Changeopening onlyInserted layers, learned prompts, bias terms alone and a compressed frozen base are the other cheap methods, and each puts the change somewhere different at a different cost.
  7. 07Keeping the Correction Separate on Purposeopening onlyNot folding the correction into the weights lets one copy of a base model serve hundreds of different adaptations at once, which changes what a per-customer model costs.
  8. 08Deciding What Actually Needs Changingopening onlyMost gaps that look like they need training turn out to need a better prompt or a lookup, and knowing which kind of gap you have saves the whole expense.