Chance, Spread and What to Expect
A List, and a Number on Each Item
A distribution is nothing more than a list of the things that could happen with a weight on each, and the two rules the weights obey decide everything that follows.
Start with something that is going to happen but has not happened yet. A coin
about to be flipped, a word about to be generated, a request about to arrive at a
server. The first thing to write down is not a probability at all. It is a list
of the things that could happen.
That list has to satisfy two conditions, and almost every muddle in the subject
is a failure of one of them. It has to be complete: when the thing happens, what
happened must be on the list. And its items must not overlap: exactly one of them
can be what happened, not two at once. A list of the possible outcomes of a coin
flip is heads and tails. A list consisting of heads, tails, and an outcome
described as a high flip is not a list at all, because two of those can be true
together.
Only once there is a list do the numbers go on it.
- an index running over the items of the list
- the weight attached to the i-th outcome
- how many outcomes the list has
- add up what follows, over every item of the list
The second rule is doing more work than it looks. It says the list is complete,
in the language of the numbers rather than the language of the list: if the
weights add to less than one then something that could happen has been left off,
and if they add to more than one then something is being counted twice. Whenever
a set of weights refuses to add to one, the fault is almost always in the list
rather than in the arithmetic.
Events, which are just groups
Rarely is the question about a single outcome. It is about a group of them: will
the request fail in any way at all, will the generated word be a noun, will the
reward be positive. A group of outcomes taken together is called an event, and
its probability is the sum of the weights of the outcomes in it.
- add the weight of every outcome that belongs to the group
- the probability of everything outside the group
- which is one minus the group's own probability, since the whole list sums to one
On the four outcome list above, the event that the request did not plainly
succeed is the last three items, with total weight six hundredths. Reading it the
other way, as one minus the weight of the first item, gives the same answer with
less arithmetic, and that shortcut gets more valuable the longer the list is.
| first | second | third | legal | |
|---|---|---|---|---|
| a legal one | 0.2 | 0.3 | 0.5 | 1.0 |
| sums to nine tenths | 0.2 | 0.3 | 0.4 | 0.0 |
| has a negative weight | 0.7 | 0.5 | -0.2 | 0.0 |
| all the weight on one ou | 0.0 | 0.0 | 1.0 | 1.0 |
Where the numbers come from
This is the part that textbooks hurry past and that matters most in practice.
There are exactly two ways to obtain a set of weights.
The first is to count. Run the thing many times, count how often each outcome
occurred, and divide by the number of runs. This is what a measured failure rate
is, what a word frequency is, and what nearly every number in a machine learning
system ultimately rests on. It requires that the thing has happened many times
already, and it requires a willingness to say that those past occasions and the
future one are the same kind of event, which is a judgement and not a
calculation.
The second is to assume. Nobody has flipped this particular coin ten thousand
times. The weight of a half comes from a claim about the symmetry of the coin
itself. This is faster, applies to things that have never happened, and is only
as good as the claim behind it.
Neither source is more respectable than the other, but they fail differently. A
counted weight fails when the conditions change, so that the past occasions were
not the same kind of event as the future one. An assumed weight fails when the
assumption is simply wrong. The error worth naming is treating an assumption as
though it had been counted, which is what happens whenever a number arrived at by
reasoning is later quoted as though it had been measured.
Equal weights is a claim
The most common assumption is that every outcome on the list gets the same
weight. It feels like refusing to assume anything, a kind of neutrality, and it
is nothing of the sort.
The reason is that it depends completely on how the list was cut up. Consider a
part that either works or fails. Two outcomes, so equal weights give a half each.
Now notice that failure comes in two kinds, a clean stop and a silent corruption,
and write the list with three items instead. Equal weights now give a third to
working and two thirds to failing. The situation has not changed at all. Only the
list has, and the supposedly neutral assumption has moved the weight on failure
from a half to two thirds.
So equal weights is a definite claim, and it is a claim about the list as much as
about the world. It is a good claim when the items of the list are genuinely
interchangeable, as the faces of a well made die are. It is a bad claim whenever
the list was written by someone thinking about the problem, because the way a
person chops a situation into cases carries their sense of what matters, and the
uniform distribution then quietly encodes that sense as though it were knowledge.
What survives all of this is small and solid. A list. A weight on each item, not
negative, adding to one. Groups of items, whose weight is a sum. That is the
entire apparatus, and every remaining lesson in this course is built from it
without adding anything new.
Recap
- A distribution is a list of outcomes with a number attached to each; the numbers may not be negative and they must add to one, and nothing else is required of them.
- An event is a group of outcomes, and its probability is the sum of their weights, which is why every rule about events turns out to be a rule about adding.
- The weights have to come from somewhere, either counted from things that happened or assumed, and treating an assumption as a measurement is the most common error in the subject.
This is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents