ContentsThe library

How a Filter Reads a Picture

Throwing Away Where, to Keep What

Last timeHow Far a Unit Can See

Reducing the resolution is the cheapest way to widen a view and the only way to make a network tolerate small shifts, and it is paid for in location.

Every layer so far has had weights. Pooling has none. Divide a map into blocks of

two by two, and replace each block with a single number: the largest of the four,

or their average. That is the whole operation. The map comes out half the size on

each side, a quarter of the area, and nothing was learned to achieve it.

FIG 1The two kinds of pooling
take the largest of the four, which keeps the strongest response and discards the rest
take the average instead, which keeps a blurred version of all four
The depth is untouched: pooling works on each channel independently and never mixes them. The maximum is the usual choice, on the reasoning that a map of responses should report the strongest evidence in a neighbourhood rather than dilute it with the three places the pattern was absent.

Because there is nothing to learn, there is nothing to tune and nothing to go

wrong. The gradient is simple too: for the maximum, the whole gradient goes to

whichever of the four inputs was largest and the other three get nothing, which

is a routing decision rather than a computation.

What the reduction buys

Two things, and they are different. The first is cost, and it is the reason

pooling appeared at all. A quarter as many positions means a quarter of the

arithmetic in every layer that follows, and it means the reach of those layers

grows twice as fast in input pixels, by the argument of the previous lesson.

The second is subtler and was the original justification. Consider a pattern

sitting at an even position, and pool over blocks of two. Move the pattern one

pixel to the right and it is still in the same block, so the maximum is the same

and the output is unchanged. The layer has become insensitive to a movement

that the convolution beneath it faithfully tracked.

The lesson stops here

4 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01The Same Few Numbers, Everywhere
  2. 02The Constraint That Pays for Itselfopening only
  3. 03Where the Window Fits, and How Often It Stopsopening only
  4. 04One Filter Is Never Enoughopening only
  5. 05The Slow Widening of the Viewopening only
  6. 06Throwing Away Where, to Keep Whatyou are here
  7. 07Two Discounts, Each With a Conditionopening only
  8. 08It Was a Matrix Multiply All Alongopening only

Read alongside