Two Discounts, Each With a Condition
Last timeDiscarding Positions on Purpose
Spacing the taps of a window apart buys reach for nothing. Splitting a filter into two cheaper ones buys width for a ninth. Both come with something given up.
The previous two lessons left a tension. Reach grows far too slowly at full
resolution, and the cure, reducing the resolution, destroys the location that
some tasks need. The first half of this lesson is a way out of that tension that
costs nothing, which is rare enough to be worth looking at closely.
Take the nine taps of a three by three window and, instead of reading nine
adjacent positions, read nine positions spaced two apart. The window now covers a
five by five region. The weights are still nine, the multiplications are still
nine, and the output is still at full resolution.
- the spacing, one for an ordinary window and two or more for a spaced one
- the width of the region covered, which grows in proportion to the spacing
- the same nine weights as before, unchanged in number
- the only modification: the offsets are multiplied by the spacing
| spacing | region covered | cumulative reach | weights used | |
|---|---|---|---|---|
| first | 1 | 3 | 3 | 9 |
| second | 2 | 5 | 7 | 9 |
| third | 4 | 9 | 15 | 9 |
| fourth | 8 | 17 | 31 | 9 |
The reach of such a stack grows geometrically, exactly as it did under resolution
reductions in the earlier lesson, and for the same reason: the contribution of
each layer is multiplied by how far apart the positions it reads are. The
difference is that resolution reductions achieve that spacing by throwing
positions away and spaced windows achieve it by skipping over positions that are
still there.
The holes are not free after all
A window with a spacing of two reads positions at even offsets only. Stack two
such layers and an output at an even position depends only on even positions of
the input. Stack several and the map separates into independent lattices, each
developing its own idea of the picture with no communication between them.
The symptom is a visible chequering in the output, which in a segmentation map
looks like a fine grid laid over the labels. The cause is structural rather than
a training failure, so no amount of data removes it.
The standard remedy is to vary the spacing through the stack rather than doubling
it uniformly. Spacings of one, two and three in succession have no common factor,
so every offset is eventually combined with every other and the lattices merge.
This is the sort of fix that looks arbitrary until the cause is understood and
obvious afterwards.
The lesson stops here
2 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents