ContentsThe library

Adapting a Model You Did Not Train

Keeping the Correction Separate on Purpose

Last timeThe Other Ways to Spend Almost Nothing

Not folding the correction into the weights lets one copy of a base model serve hundreds of different adaptations at once, which changes what a per-customer model costs.

The previous lessons treated folding the correction into the weights as the

point of the method. There is one situation where you deliberately do not fold,

and it is the situation that most changes what adaptation is for.

Declining to fold

Folding produces one adapted model that costs exactly what the original cost.

That is what you want when there is one adaptation. A service with four hundred

customers, each wanting the model to behave slightly differently, wants

something else.

FIG 1One request, one base, one of many corrections
The base is shared and the correction is swapped. Nothing about this requires retraining or redeployment to add a customer, which is the property the whole arrangement exists to provide.
FIG 2Memory to serve k adaptations from one machine
one full copy of the base model, held once
how many adaptations are resident
the values in one correction, across all corrected matrices
bytes per value
The first term is fixed and the second is tiny per adaptation. The arrangement that this replaces has the first term multiplied by k instead, which is the entire argument.
FIG 3Memory against the number of adaptations served
0.00625.001250.001875.002500.001.04.88.512.316.0adaptations served from one machine
one shared base, corrections kept separatea folded copy of the model for each
The lower line is almost flat because an adaptation is about a thousandth of the size of the model it corrects. The upper line is the arrangement folding forces you into, and it runs off the top of any plausible machine almost immediately.
FIG 4Memory for the corrections alone, in gigabytes
bottleneck 8bottleneck 16bottleneck 64
1 adaptation0.020.030.13
10 adaptations0.170.341.34
100 adaptations1.683.3613.42
1000 adaptations16.7833.55134.22
The base model is not in this table, only the corrections. The marked cell is a hundred customers at a sensible bottleneck width, costing under two gigabytes on top of a base that costs a hundred and forty.

What the memory looks like

It is worth being clear that the decision is made at deployment and not at

training. The same trained correction can be folded into a copy of the weights

for one customer who wants a dedicated model, and kept separate on a shared

machine for everybody else. Nothing about the training run has to know which

arrangement it is destined for, which means the choice can be revisited later

without retraining anything.

The lesson stops here

5 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01Almost All of the Work Is Already Done
  2. 02The Obvious Approach and What It Costsopening only
  3. 03The Abilities Nobody Was Trying to Changeopening only
  4. 04Writing the Correction as Two Thin Sheetsopening only
  5. 05Spending a Fixed Budget in the Right Placesopening only
  6. 06Four Other Places to Put the Changeopening only
  7. 07Keeping the Correction Separate on Purposeyou are here
  8. 08Deciding What Actually Needs Changingopening only

Read alongside