ContentsThe library

How a Filter Reads a Picture

The Slow Widening of the View

Last timeMany Filters, Stacked

A unit can only be influenced by the input positions that feed it. That region grows by two pixels a layer, which is far too slow until stride enters the arithmetic.

Pick one number in the output of the twentieth layer and ask which pixels of the

original photograph could possibly change it. Trace the dependency backwards. In

the last layer it reads nine positions of the layer below; each of those read

nine of the layer below that, overlapping heavily. Keep going and the set of

original pixels involved is some square patch. Everything outside that patch is

irrelevant to this number by construction: no training, no amount of data, no

clever weighting can make it matter.

That patch is the bound on what this unit can know. It is the reason a unit early

in a network cannot be a face detector even in principle, and the reason the

design of a convolutional network is largely an argument about how fast to widen

it.

FIG 1Reach through a stack at full resolution
how many input pixels across the region is, after L layers
what layer l adds, which is two for a three-wide window
the total extension, linear in depth
Linear growth is the problem. Twenty three-wide layers reach forty one pixels. A network that must decide whether a photograph of two hundred and twenty four pixels a side contains a dog needs units that see most of it, and at two pixels a layer that would take a hundred and twelve layers of full resolution convolution, at a cost nobody would pay.
FIG 2Two layers reaching five
Two three-wide layers reach exactly as far as one five-wide layer. The difference is that the pair has a nonlinearity in the middle and fewer weights, and that the influence of the five inputs is uneven rather than flat.

Where the reach actually comes from

The linear growth above assumed every layer works at full resolution. As soon as

a layer has a stride of two, the positions its output describes are two input

pixels apart, so a three-wide window in the next layer no longer spans three

input pixels but five, and the layer after that spans more again.

FIG 3The recurrence, with strides
the reach after l layers, in input pixels
the window width of layer l
the spacing of the positions layer l reads, being the product of all earlier strides
the stride of layer i, one for an ordinary layer
The product is what changes the character of the growth. Every stride of two doubles the contribution of every layer that follows, so reach grows geometrically in the number of reductions and only linearly in the number of layers between them.
FIG 4Six layers, a halving at every second one
layerstride of this layerspacing of its inputsreach in input pixels
first1113
second2215
third3129
fourth42213
fifth51421
sixth62429
Six three-wide layers reach twenty nine pixels here, against thirteen if all six had been at full resolution. The gap widens with every further reduction, and by the fifth halving a single unit sees most of the picture. The fourth column is doing the work the third column makes possible.
FIG 5Reach against depth, with and without reductions
0.0075.00150.00225.00300.001.03.86.59.312.0number of three-wide layers
every layer at full resolutiona stride of two at every second layer
The two curves part company almost immediately. At twelve layers one has reached twenty five pixels and the other two hundred and fifty three, which is more than a standard input is wide. Everything about how deep networks are arranged into stages of falling resolution follows from this picture.

Small windows, stacked

The diagram above contained the second argument of the lesson. Two three-wide

layers reach five input pixels, which is exactly what one five-wide layer

reaches. So the choice between them is not about reach at all.

Count the weights, per pair of channels. The five-wide window has twenty five.

The two three-wide windows have nine each, eighteen in total, which is a quarter

less. The arithmetic follows the same ratio. And between the two three-wide

layers there is a nonlinearity, so the pair computes a strictly larger class of

functions than the single layer, which is one linear map followed by one

nonlinearity however wide its window.

Three advantages, no disadvantage. This is the argument that settled window

sizes at three almost everywhere, and the same reasoning extends: three

three-wide layers reach seven pixels with twenty seven weights against forty nine

for a seven-wide window, and so on with the gap widening.

The lesson stops here

1 more paragraph to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01The Same Few Numbers, Everywhere
  2. 02The Constraint That Pays for Itselfopening only
  3. 03Where the Window Fits, and How Often It Stopsopening only
  4. 04One Filter Is Never Enoughopening only
  5. 05The Slow Widening of the Viewyou are here
  6. 06Throwing Away Where, to Keep Whatopening only
  7. 07Two Discounts, Each With a Conditionopening only
  8. 08It Was a Matrix Multiply All Alongopening only

Read alongside