The Numbers A Machine Cannot Hold
A number format buys you a range of magnitudes and a number of significant digits, and both budgets are small enough that ordinary arithmetic walks off either end of them.
Everything in this course follows from a limitation that is easy to state and easy
to forget. A machine does not hold numbers. It holds a fixed number of bits, and it
uses them to name a finite set of values that stands in for the real numbers. The
substitution is excellent over a middle band and fails at both edges, and a deep
network is a machine for pushing values towards those edges.
The arrangement is the same one used in scientific notation. A sign, a handful of
significant digits, and a power to shift them by. The digits and the power come out
of separate fields in the format, which is why range and precision are two separate
budgets and why they can be traded against each other.
- one bit of sign
- the significant digits, which set how finely nearby values can be told apart
- the stored exponent, which sets the magnitude the digits are placed at
- a fixed offset so that the exponent can be negative as well as positive
The two ends, both of them quiet
Multiply enough numbers greater than one together and you pass the largest value
the exponent can name. The result becomes infinity. Nothing stops. The next
operation involving that infinity may produce a value that is not a number at all,
and that value propagates through every sum it participates in, so a single
overflow in one layer can turn an entire batch of predictions into nothing.
Go the other way and the failure is gentler and more insidious. Below a certain
magnitude the format runs out of exponent and starts borrowing from the digits, so
precision degrades before the value finally rounds to zero. There is no boundary
you cross with a bang. A quantity simply becomes less and less accurate and then
becomes nothing, and a gradient that has become nothing cannot be distinguished
from a parameter that is already optimal.
Precision is relative, which changes what you can ask for
The representable values are not evenly spaced. They are spaced in proportion to
their own size, because the digits are always placed at whatever magnitude the
exponent selects. Near one the spacing is tiny. Near a million it is a million
times larger. The practical reading is that a format promises you a fixed number of
significant digits, never a fixed absolute accuracy.
This is the first real argument for keeping activations near one. It is not an
aesthetic preference. If a layer's outputs drift up to the thousands, you still have
the same seven significant digits, but the smallest difference the format can
express between two of those outputs has grown by a factor of a thousand, and
updates smaller than that gap simply do not happen.
Two ways arithmetic destroys digits
There are two patterns worth recognising on sight, because both appear inside
ordinary layer computations.
The first is subtracting two numbers that are nearly equal. The leading digits
agree and cancel, and what survives is the trailing digits, which were the least
reliable part of each input. A difference computed this way can be entirely
rounding error while looking perfectly plausible. This is why a variance is
computed by accumulating deviations from the mean rather than by subtracting the
square of the mean from the mean of the squares, even though the second formula is
algebraically identical and shorter.
The second is adding a small number to a much larger one. If the small number falls
below the gap between neighbouring representable values at the large number's
magnitude, the sum is simply the large number. The addition had no effect. Repeat
this across a long summation and most of the terms are discarded, which is exactly
the shape of a dot product in a wide layer.
| step | step | running total | what the format did |
|---|---|---|---|
| 1 | start | 1.000 | the total is held with about three signi |
| 2 | add 0.0004 | 1.000 | the term is smaller than the gap at this |
| 3 | repeat 999 times | 1.000 | every single term is discarded for the s |
| 4 | true answer | 1.400 | the terms were real and should have move |
| 5 | same sum, wider format | 1.400 | nothing about the mathematics changed, o |
What practice settles on
Three formats cover most of what you will meet. Full precision gives roughly seven
significant digits across a very wide range of magnitudes and is the historical
default. Half precision uses half the storage, which halves the memory traffic and
roughly doubles the arithmetic throughput, but keeps only about three digits and
compresses the range dramatically. The third option keeps the wide range of full
precision and spends the saving entirely on digits instead, leaving only about three
of them.
That third trade reads as strange until you hold the two failure modes side by
side. Too few digits makes updates imprecise, which training tolerates well,
because the next step corrects it. Too little range makes quantities become zero or
infinity, which training does not tolerate at all, because there is no recovering
from a gradient of nothing. Given a fixed budget, range is the one to protect.
| bits per number | decimal digits held | orders of magnitude of r | |
|---|---|---|---|
| full precision | 32.0 | 7.2 | 76.0 |
| half precision | 16.0 | 3.3 | 8.0 |
| wide-range half | 16.0 | 2.4 | 76.0 |
| eight bit | 8.0 | 1.2 | 4.0 |
So the ground rule for everything that follows is that quantities inside a network
should stay near one, not because one is special, but because that is where a
format's digits are worth the most and where neither end of the range is close. The
next lesson asks what a stack of layers does to a quantity that starts there, and
the answer is that it does not stay.
Recap
- Range and precision are separate budgets spent from separate parts of the format, and a format can be generous with one while being miserly with the other.
- Precision is relative, so the gap between neighbouring representable numbers grows in proportion to the size of the number, and absolute accuracy is worst where numbers are largest.
- Both ends of the range fail silently: too large becomes infinity, too small becomes zero, and arithmetic continues afterwards as though nothing happened.
This is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents