ContentsThe library

Making a Model Smaller

A Zero Still Occupies Its Place in the Row

Last timeRounding After, or Training With

Removing scattered weights and removing whole rows sound like the same technique. Only one of them makes a model smaller or faster on ordinary hardware.

Removing weights sounds like the same kind of idea as storing them in fewer

bits, and it is not. Fewer bits keeps every weight and describes each one more

crudely. Removal keeps some weights exactly as they were and deletes others.

The important split is inside removal itself, and it is the thing this lesson

exists for. There are two families, they are reported with the same vocabulary

and the same impressive percentages, and one of them does nothing at all to

the memory or the speed of an ordinary served model.

What scattered removal actually produces

Take a block of weights, find the smallest ones by absolute value, and set them

to zero. The usual report is that the layer is now fifty per cent sparse.

FIG 1A small block after half its weights were removed
c1c2c3c4c5c6
row 10.410.000.000.630.000.22
row 20.000.370.580.000.190.00
row 30.260.000.000.000.440.31
row 40.000.510.330.280.000.00
row 50.180.000.470.000.000.39
row 60.000.290.000.360.550.00
Eighteen of thirty-six weights removed. Count what changed about the block: it is still six rows by six columns, it still occupies thirty-six numbers of storage, and multiplying a vector by it still performs thirty-six multiplications. The model is reported as fifty per cent sparse and is exactly the same size and speed it was.

This is the whole objection. A matrix is stored as a dense rectangle of numbers

and multiplied by hardware that works through every position in it. A zero

occupies its bytes like any other value and consumes its multiplication like

any other value. Nothing was saved, because nothing about the shape changed.

Two things can rescue it. A sparse storage format stores only the nonzero

values plus their positions, which saves memory once the sparsity is high

enough to pay for the position bookkeeping, and is usually slower to compute

with. Or hardware with explicit support for a fixed pattern, such as two

nonzeros in every group of four, which does deliver a genuine speedup but

constrains which weights you may remove.

Absent one of those, scattered removal is a research result rather than a

deployment technique, and it is worth knowing which one you are being shown.

The lesson stops here

5 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01Three Budgets, and Only One of Them Is the Weights
  2. 02Round It, Store the Integer, Multiply It Backopening only
  3. 03Most of the Model Does Not Care, and You Have to Find the Part That Doesopening only
  4. 04Round the Finished Model, or Train One That Expects Itopening only
  5. 05A Zero Still Occupies Its Place in the Rowyou are here
  6. 06The Wrong Answers Are Where the Teaching Isopening only
  7. 07The Average Moved Half a Point and the Model Cannot Add Any Moreopening only
  8. 08The Savings Multiply and So Does the Damageopening only

Read alongside