ContentsThe library

Keeping The Numbers In Range

The Numbers A Machine Cannot Hold

A number format buys you a range of magnitudes and a number of significant digits, and both budgets are small enough that ordinary arithmetic walks off either end of them.

Everything in this course follows from a limitation that is easy to state and easy

to forget. A machine does not hold numbers. It holds a fixed number of bits, and it

uses them to name a finite set of values that stands in for the real numbers. The

substitution is excellent over a middle band and fails at both edges, and a deep

network is a machine for pushing values towards those edges.

The arrangement is the same one used in scientific notation. A sign, a handful of

significant digits, and a power to shift them by. The digits and the power come out

of separate fields in the format, which is why range and precision are two separate

budgets and why they can be traded against each other.

FIG 1What a stored number actually is
one bit of sign
the significant digits, which set how finely nearby values can be told apart
the stored exponent, which sets the magnitude the digits are placed at
a fixed offset so that the exponent can be negative as well as positive
The two interesting fields are independent. Spend bits on the digits and you can distinguish values that are very close together. Spend them on the exponent and you can name very large and very small magnitudes. A format of a given size must choose, and the choice has consequences that show up only in deep networks.

The two ends, both of them quiet

Multiply enough numbers greater than one together and you pass the largest value

the exponent can name. The result becomes infinity. Nothing stops. The next

operation involving that infinity may produce a value that is not a number at all,

and that value propagates through every sum it participates in, so a single

overflow in one layer can turn an entire batch of predictions into nothing.

Go the other way and the failure is gentler and more insidious. Below a certain

magnitude the format runs out of exponent and starts borrowing from the digits, so

precision degrades before the value finally rounds to zero. There is no boundary

you cross with a bang. A quantity simply becomes less and less accurate and then

becomes nothing, and a gradient that has become nothing cannot be distinguished

from a parameter that is already optimal.

FIG 2The band a format can name
the whole working band of half precisionthe working band of full precision-40-202040-38smallest ordinary value in full precision-8smallest value half precision can hold accurately0one, where the digits are most useful5largest value half precision can hold38largest value in full precisionpower of ten of the magnitude
Note how narrow the inner band is. Half precision covers about thirteen powers of ten in total, which sounds ample until you remember that a product of fifty layer outputs can move through that many powers on its own. The wide band is not a luxury, it is what stops the intermediate quantities of a deep network from leaving the format.

Precision is relative, which changes what you can ask for

The representable values are not evenly spaced. They are spaced in proportion to

their own size, because the digits are always placed at whatever magnitude the

exponent selects. Near one the spacing is tiny. Near a million it is a million

times larger. The practical reading is that a format promises you a fixed number of

significant digits, never a fixed absolute accuracy.

This is the first real argument for keeping activations near one. It is not an

aesthetic preference. If a layer's outputs drift up to the thousands, you still have

the same seven significant digits, but the smallest difference the format can

express between two of those outputs has grown by a factor of a thousand, and

updates smaller than that gap simply do not happen.

FIG 3How far apart neighbouring values are
0.000.300.600.901.201.0250.8500.5750.31000.0magnitude of the number
a format with more significant digitsa format with fewer significant digits
Both lines pass through the origin, which is the whole point: the gap is proportional to the magnitude, so the accuracy you actually get depends on where your numbers have drifted to. The steeper line is the cheaper format. Halving the storage does not halve the accuracy, it multiplies the gaps by a much larger factor, because digits are being removed rather than scaled.

Two ways arithmetic destroys digits

There are two patterns worth recognising on sight, because both appear inside

ordinary layer computations.

The first is subtracting two numbers that are nearly equal. The leading digits

agree and cancel, and what survives is the trailing digits, which were the least

reliable part of each input. A difference computed this way can be entirely

rounding error while looking perfectly plausible. This is why a variance is

computed by accumulating deviations from the mean rather than by subtracting the

square of the mean from the mean of the squares, even though the second formula is

algebraically identical and shorter.

The second is adding a small number to a much larger one. If the small number falls

below the gap between neighbouring representable values at the large number's

magnitude, the sum is simply the large number. The addition had no effect. Repeat

this across a long summation and most of the terms are discarded, which is exactly

the shape of a dot product in a wide layer.

FIG 4A thousand small additions, in a narrow format
stepsteprunning totalwhat the format did
1start1.000the total is held with about three signi
2add 0.00041.000the term is smaller than the gap at this
3repeat 999 times1.000every single term is discarded for the s
4true answer1.400the terms were real and should have move
5same sum, wider format1.400nothing about the mathematics changed, o
5 steps
The sum is not slightly wrong. It is completely wrong, and no step in it raised a complaint. This is the failure that mixed precision training has to design around, and the usual answer is to accumulate in a wider format even when the multiplications are done in a narrow one.

What practice settles on

Three formats cover most of what you will meet. Full precision gives roughly seven

significant digits across a very wide range of magnitudes and is the historical

default. Half precision uses half the storage, which halves the memory traffic and

roughly doubles the arithmetic throughput, but keeps only about three digits and

compresses the range dramatically. The third option keeps the wide range of full

precision and spends the saving entirely on digits instead, leaving only about three

of them.

That third trade reads as strange until you hold the two failure modes side by

side. Too few digits makes updates imprecise, which training tolerates well,

because the next step corrects it. Too little range makes quantities become zero or

infinity, which training does not tolerate at all, because there is no recovering

from a gradient of nothing. Given a fixed budget, range is the one to protect.

FIG 5Four formats, in numbers
bits per numberdecimal digits heldorders of magnitude of r
full precision32.07.276.0
half precision16.03.38.0
wide-range half16.02.476.0
eight bit8.01.24.0
Read the second and third columns against each other and the trade is visible. The marked cell is the one that decided the matter: keeping the full range at sixteen bits, at the cost of one decimal digit of precision, made large-scale training routine, because the narrow-range format needed a separate scaling trick to stop small values rounding away and the wide-range one did not. Eight bit formats are now common for running a trained model and still delicate for training one.

So the ground rule for everything that follows is that quantities inside a network

should stay near one, not because one is special, but because that is where a

format's digits are worth the most and where neither end of the range is close. The

next lesson asks what a stack of layers does to a quantity that starts there, and

the answer is that it does not stay.

Recap

  • Range and precision are separate budgets spent from separate parts of the format, and a format can be generous with one while being miserly with the other.
  • Precision is relative, so the gap between neighbouring representable numbers grows in proportion to the size of the number, and absolute accuracy is worst where numbers are largest.
  • Both ends of the range fail silently: too large becomes infinity, too small becomes zero, and arithmetic continues afterwards as though nothing happened.

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

NextFifty Multiplications Later →

The rest of this course

  1. 01The Numbers A Machine Cannot Holdyou are here
  2. 02The Compounding Nobody Budgets Foropening only
  3. 03The Decision Made Before Anything Runsopening only
  4. 04Putting The Numbers Back Where They Belongopening only
  5. 05Two Decisions That Look Like Detailsopening only
  6. 06The Road That Goes Roundopening only
  7. 07The Same Trouble, Running Backwardsopening only
  8. 08What A Training Curve Is Telling Youopening only

Read alongside