The Comparison That Actually Means Something
Last timeAll of It in Memory, a Little of It in Use
Sparsity is a third dial alongside size and data, and whether turning it up helps depends on what you compare against, what resource is scarce, and how many requests arrive at once.
Everything so far has been about how the arrangement works and what goes wrong
with it. This lesson is about whether to use it, which is a different question
and one that is usually answered badly.
Sparsity is a dial, not a decision
It helps to stop thinking of sparse and dense as two kinds of model. There is
one model with a setting: what share of its parameters sit idle on any given
input. Zero is the ordinary dense model. Set it to seven eighths and a token
touches one copy in eight.
Once it is a setting, the question is what value to use, and that has been
measured rather than argued about. With a fixed training budget you can buy a
smaller dense model trained on more data, or a much larger sparse model that
performs the same arithmetic per token. The answer is not constant.
The mechanism is not mysterious. Quality improves with parameters and with data,
and sparsity buys parameters without buying the arithmetic to run them. When
arithmetic is the binding constraint, which it increasingly is at large budgets,
that purchase is a good one. When data or memory is binding instead, it is not.
The lesson stops here
7 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents