The Constraint That Pays for Itself
Last timeOne Small Window, Moved Everywhere
Sharing one small window across every position cuts the parameters by orders of magnitude, but the saving is the lesser half of what the constraint actually buys.
Take a colour photograph of two hundred and fifty six pixels on a side. That is
sixty five thousand positions and, with three colours, about two hundred thousand
numbers. Build the most ordinary layer there is, one that connects every input to
every output, and give it as many outputs as there were inputs. The weight count
is two hundred thousand squared, which is about forty billion. There is no
machine on which that is a sensible first layer, and no training set large enough
to pin down that many numbers.
Now take the alternative from the previous lesson. One three by three window with
nine weights, slid over every position. The layer still produces an output at
every position, and it still reads the picture everywhere. It has nine weights.
| side length | connected layer weights | shared filter weights | |
|---|---|---|---|
| 32 pixels a side | 32 | 1048576 | 9 |
| 64 pixels a side | 64 | 16777216 | 9 |
| 128 pixels a side | 128 | 268435456 | 9 |
| 256 pixels a side | 256 | 4294967296 | 9 |
- the height and width of the window, almost always three
- how many channels come in, the subject of a later lesson
- how many filters this layer has, each producing one output channel
- appearing a second time as one bias per filter
The absence of the picture size from that formula is worth pausing on. It means
the layer is not really a function of an image of a particular size. It is a rule
for turning a patch into a number, and the image size only decides how many times
the rule is applied. Nothing has to be retrained to run it on a larger photograph.
The saving is the lesser half
It is easy to read all of this as a compromise: the full layer is what one would
want, it is unaffordable, so the weights are shared and something is given up.
That reading predicts that convolution would be abandoned as memory got cheaper.
Memory has become enormously cheaper and it has not been abandoned, because the
reading is wrong.
What sharing encodes is a claim about photographs. A vertical edge is a vertical
edge whether it is in the top left or the bottom right. The thing that makes it
an edge is the local arrangement of brightness, and that arrangement means the
same thing everywhere. A connected layer does not know this. Shown ten thousand
edges in ten thousand places, it has to learn each one separately, and it will
cheerfully learn an edge detector for the top left corner that fails in the top
right because the training set happened not to contain one there.
The lesson stops here
2 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents