ContentsThe library

How a Filter Reads a Picture

The Same Few Numbers, Everywhere

One output value is a weighted sum of a small patch of input. Slide that same patch of weights across every position and the result is a map of where the pattern was found.

Start with a row of numbers rather than a picture, because the picture adds a

second index and nothing else. Say the row is a scan line of brightness values,

and say we have three weights. Lay the three weights over the first three values

of the row, multiply each weight by the value beneath it, add the three products,

and write the answer down. That single number is the output at the first

position. Then slide the three weights one place to the right and do it again.

FIG 1One output position, in one dimension
the input row, a long list of numbers
the filter, a short list of weights, here of length k
the output at position i, a single number
the width of the window, typically three or five and almost never large
This is a dot product between the filter and the k values of the input starting at position i. There is nothing in it but multiplication and addition, and no reference anywhere to where i is in the row, which is the point of the next lesson.

Notice what the formula does not contain. There is no term that depends on i

except through which values of the input are read. The weights are the same at

every position; only the patch beneath them changes. Run i over every valid

position and the collection of answers is as long as the row, give or take the

edges, and ordered the same way.

The output is another picture

That ordering matters more than it first seems. A layer that takes the whole row

and produces ten numbers has thrown away the geography: the ten numbers describe

the row as a whole, and there is no longer any sense in which one of them is on

the left. A convolution does the opposite. The value it writes at position i came

only from the input near position i, and it is stored in a list whose position i

means the same place.

So the output is a map. If the weights happen to describe a dark-to-light

transition, the output is bright exactly where the input had a dark-to-light

transition and near zero everywhere else. The name usually given to this is a

feature map, and the name is accurate: it is a map, in the cartographic sense, of

where a particular feature was found.

FIG 2A three-wide filter answering four patches
first valuesecond valuethird valuefilter response
flat dark1010100
an edge beginning10109080
an edge ending109090160
flat bright9090900
The filter here is minus one, zero, plus one. Flat patches give zero no matter how bright they are, because the weights sum to zero. Only the two patches containing a change produce anything, and the larger change produces the larger number. The same three weights did all four of these without being told which patch it was looking at.

The filter that produced those four numbers is a comparison: subtract what is on

the left, add what is on the right. Any patch where left and right agree gives

zero, which is why both flat rows vanish regardless of their brightness. That is

a useful thing for a first layer to do, and it is also an honest preview of the

whole method, because nothing in the arithmetic knew the word edge. The weights

were a pattern and the dot product measured agreement with it.

FIG 3The same thing with two indices
the output at one row and column
the two sums, running over the height and the width of the window
one filter weight, named by its offsets inside the window
the input value that weight lands on
Adding the second dimension changes one sum into two and nothing else. A three by three filter has nine weights, lies over nine input values at a time, and produces one number from them. Every later lesson in this course modifies the index arithmetic inside this expression rather than its shape.

Why a dot product detects anything

A dot product between two lists is large when they are large and point the same

way, and near zero when they disagree. The filter is a list, so it is a direction

in the space of small patches. The response at a position asks how much of the

patch beneath the window lies along that direction.

This is why a filter is sometimes described as a template and the operation as

template matching. The description is right for a single filter and becomes

misleading in deeper layers, where the thing being matched is not in the picture

at all but in the output of earlier filters. The arithmetic is identical; only the

input has changed.

FIG 4The whole operation, written out
python
for r in range(H - kh + 1):           # every valid row position
    for c in range(W - kw + 1):       # every valid column position
        total = 0.0
        for u in range(kh):           # the window, same weights every time
            for v in range(kw):
                total += w[u][v] * x[r + u][c + v]
        y[r][c] = total
Four loops and one multiply-accumulate. The outer two walk the output; the inner two are the dot product. The thing worth staring at is that w is read inside the inner loops and never indexed by r or c, which is the entire difference between this and a connected layer.

The two properties worth naming

Two facts about this operation get used repeatedly later, and both are visible in

the formula. First, it is linear: double the input and the output doubles, add

two inputs and the outputs add. That means the whole thing is a matrix

multiplication in disguise, a claim the final lesson makes explicit.

Second, it commutes with shifting. Move the input two places to the right and the

output moves two places to the right, with every value unchanged. The usual word

for this is equivariance, and it is the precise sense in which a convolution

treats all positions alike. A connected layer has no such property: shifting its

input changes every output, because every input position has its own weights.

FIG 5Sliding three weights along one row
stepwindow positionvalues under the windowresponse
1010, 10, 100
2110, 10, 9080
3210, 90, 9080
4390, 90, 900
5490, 90, 20-70
5 steps
One pass of the same minus one, zero, plus one filter along a single row. The two non-zero stretches mark the two places where the brightness changed, and the sign records which way. Nothing was stored between positions, so these five numbers could all have been computed at once, which is why this operation suits hardware that does thousands of multiplications in parallel.

What has and has not been decided

So far there is one window, one set of weights, and one map coming out. Three

things have been quietly left open, and each is a lesson of its own. The window

reads values that do not exist when it hangs over the edge of the picture, and

something has to be done about that. It has been slid one place at a time, which

is a choice rather than a necessity. And a single filter detects a single

pattern, where any real use needs dozens at once.

The important thing to carry forward is smaller than any of those. The operation

is a dot product repeated at every position with the same weights, and its output

is a map in the same coordinates as its input. Everything else in this course is

bookkeeping around those two sentences.

Recap

  • A convolution output at one position is nothing but a dot product between a small window of weights and the equally small patch of input sitting under it, so the whole operation is built from the most ordinary arithmetic there is.
  • Because the same weights are used at every position, the output is not a summary but a map: it has positions of its own, and the value at each one says how strongly the pattern was present at the matching place in the input.
  • The operation is linear and it commutes with shifting the input, which together are the two properties that every later lesson in this course leans on.

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

NextWhy the Same Weights Everywhere →

The rest of this course

  1. 01The Same Few Numbers, Everywhereyou are here
  2. 02The Constraint That Pays for Itselfopening only
  3. 03Where the Window Fits, and How Often It Stopsopening only
  4. 04One Filter Is Never Enoughopening only
  5. 05The Slow Widening of the Viewopening only
  6. 06Throwing Away Where, to Keep Whatopening only
  7. 07Two Discounts, Each With a Conditionopening only
  8. 08It Was a Matrix Multiply All Alongopening only

Read alongside