The Slow Widening of the View
Last timeMany Filters, Stacked
A unit can only be influenced by the input positions that feed it. That region grows by two pixels a layer, which is far too slow until stride enters the arithmetic.
Pick one number in the output of the twentieth layer and ask which pixels of the
original photograph could possibly change it. Trace the dependency backwards. In
the last layer it reads nine positions of the layer below; each of those read
nine of the layer below that, overlapping heavily. Keep going and the set of
original pixels involved is some square patch. Everything outside that patch is
irrelevant to this number by construction: no training, no amount of data, no
clever weighting can make it matter.
That patch is the bound on what this unit can know. It is the reason a unit early
in a network cannot be a face detector even in principle, and the reason the
design of a convolutional network is largely an argument about how fast to widen
it.
- how many input pixels across the region is, after L layers
- what layer l adds, which is two for a three-wide window
- the total extension, linear in depth
Where the reach actually comes from
The linear growth above assumed every layer works at full resolution. As soon as
a layer has a stride of two, the positions its output describes are two input
pixels apart, so a three-wide window in the next layer no longer spans three
input pixels but five, and the layer after that spans more again.
- the reach after l layers, in input pixels
- the window width of layer l
- the spacing of the positions layer l reads, being the product of all earlier strides
- the stride of layer i, one for an ordinary layer
| layer | stride of this layer | spacing of its inputs | reach in input pixels | |
|---|---|---|---|---|
| first | 1 | 1 | 1 | 3 |
| second | 2 | 2 | 1 | 5 |
| third | 3 | 1 | 2 | 9 |
| fourth | 4 | 2 | 2 | 13 |
| fifth | 5 | 1 | 4 | 21 |
| sixth | 6 | 2 | 4 | 29 |
Small windows, stacked
The diagram above contained the second argument of the lesson. Two three-wide
layers reach five input pixels, which is exactly what one five-wide layer
reaches. So the choice between them is not about reach at all.
Count the weights, per pair of channels. The five-wide window has twenty five.
The two three-wide windows have nine each, eighteen in total, which is a quarter
less. The arithmetic follows the same ratio. And between the two three-wide
layers there is a nonlinearity, so the pair computes a strictly larger class of
functions than the single layer, which is one linear map followed by one
nonlinearity however wide its window.
Three advantages, no disadvantage. This is the argument that settled window
sizes at three almost everywhere, and the same reasoning extends: three
three-wide layers reach seven pixels with twenty seven weights against forty nine
for a seven-wide window, and so on with the gap widening.
The lesson stops here
1 more paragraph to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents