ContentsThe library

How a Filter Reads a Picture

The Constraint That Pays for Itself

Last timeOne Small Window, Moved Everywhere

Sharing one small window across every position cuts the parameters by orders of magnitude, but the saving is the lesser half of what the constraint actually buys.

Take a colour photograph of two hundred and fifty six pixels on a side. That is

sixty five thousand positions and, with three colours, about two hundred thousand

numbers. Build the most ordinary layer there is, one that connects every input to

every output, and give it as many outputs as there were inputs. The weight count

is two hundred thousand squared, which is about forty billion. There is no

machine on which that is a sensible first layer, and no training set large enough

to pin down that many numbers.

Now take the alternative from the previous lesson. One three by three window with

nine weights, slid over every position. The layer still produces an output at

every position, and it still reads the picture everywhere. It has nine weights.

FIG 1The same layer, two ways
side lengthconnected layer weightsshared filter weights
32 pixels a side3210485769
64 pixels a side64167772169
128 pixels a side1282684354569
256 pixels a side25642949672969
The middle column is one weight per input value per output unit, both counted in greyscale pixels. It grows with the fourth power of the side length, so doubling the resolution multiplies it by sixteen. The right column does not move. The marked entry is the whole argument in one number.
FIG 2What a convolution layer actually has to learn
the height and width of the window, almost always three
how many channels come in, the subject of a later lesson
how many filters this layer has, each producing one output channel
appearing a second time as one bias per filter
The picture size is absent from this expression. A layer trained on small images has exactly the same number of parameters as the same layer trained on large ones, which is why a trained convolutional network can be run on an input size it never saw.

The absence of the picture size from that formula is worth pausing on. It means

the layer is not really a function of an image of a particular size. It is a rule

for turning a patch into a number, and the image size only decides how many times

the rule is applied. Nothing has to be retrained to run it on a larger photograph.

FIG 3Weights per output unit, against picture size
0.001250.002500.003750.005000.008.022.036.050.064.0picture side length in pixels
connected layer, reading every pixelthree by three filter, reading nine
One curve grows without limit and the other is a horizontal line near the bottom of the frame. The gap is already a factor of four hundred at sixty four pixels a side, and real photographs are many times larger than that.

The saving is the lesser half

It is easy to read all of this as a compromise: the full layer is what one would

want, it is unaffordable, so the weights are shared and something is given up.

That reading predicts that convolution would be abandoned as memory got cheaper.

Memory has become enormously cheaper and it has not been abandoned, because the

reading is wrong.

What sharing encodes is a claim about photographs. A vertical edge is a vertical

edge whether it is in the top left or the bottom right. The thing that makes it

an edge is the local arrangement of brightness, and that arrangement means the

same thing everywhere. A connected layer does not know this. Shown ten thousand

edges in ten thousand places, it has to learn each one separately, and it will

cheerfully learn an edge detector for the top left corner that fails in the top

right because the training set happened not to contain one there.

The lesson stops here

2 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01The Same Few Numbers, Everywhere
  2. 02The Constraint That Pays for Itselfyou are here
  3. 03Where the Window Fits, and How Often It Stopsopening only
  4. 04One Filter Is Never Enoughopening only
  5. 05The Slow Widening of the Viewopening only
  6. 06Throwing Away Where, to Keep Whatopening only
  7. 07Two Discounts, Each With a Conditionopening only
  8. 08It Was a Matrix Multiply All Alongopening only

Read alongside