The Same Few Numbers, Everywhere
One output value is a weighted sum of a small patch of input. Slide that same patch of weights across every position and the result is a map of where the pattern was found.
Start with a row of numbers rather than a picture, because the picture adds a
second index and nothing else. Say the row is a scan line of brightness values,
and say we have three weights. Lay the three weights over the first three values
of the row, multiply each weight by the value beneath it, add the three products,
and write the answer down. That single number is the output at the first
position. Then slide the three weights one place to the right and do it again.
- the input row, a long list of numbers
- the filter, a short list of weights, here of length k
- the output at position i, a single number
- the width of the window, typically three or five and almost never large
Notice what the formula does not contain. There is no term that depends on i
except through which values of the input are read. The weights are the same at
every position; only the patch beneath them changes. Run i over every valid
position and the collection of answers is as long as the row, give or take the
edges, and ordered the same way.
The output is another picture
That ordering matters more than it first seems. A layer that takes the whole row
and produces ten numbers has thrown away the geography: the ten numbers describe
the row as a whole, and there is no longer any sense in which one of them is on
the left. A convolution does the opposite. The value it writes at position i came
only from the input near position i, and it is stored in a list whose position i
means the same place.
So the output is a map. If the weights happen to describe a dark-to-light
transition, the output is bright exactly where the input had a dark-to-light
transition and near zero everywhere else. The name usually given to this is a
feature map, and the name is accurate: it is a map, in the cartographic sense, of
where a particular feature was found.
| first value | second value | third value | filter response | |
|---|---|---|---|---|
| flat dark | 10 | 10 | 10 | 0 |
| an edge beginning | 10 | 10 | 90 | 80 |
| an edge ending | 10 | 90 | 90 | 160 |
| flat bright | 90 | 90 | 90 | 0 |
The filter that produced those four numbers is a comparison: subtract what is on
the left, add what is on the right. Any patch where left and right agree gives
zero, which is why both flat rows vanish regardless of their brightness. That is
a useful thing for a first layer to do, and it is also an honest preview of the
whole method, because nothing in the arithmetic knew the word edge. The weights
were a pattern and the dot product measured agreement with it.
- the output at one row and column
- the two sums, running over the height and the width of the window
- one filter weight, named by its offsets inside the window
- the input value that weight lands on
Why a dot product detects anything
A dot product between two lists is large when they are large and point the same
way, and near zero when they disagree. The filter is a list, so it is a direction
in the space of small patches. The response at a position asks how much of the
patch beneath the window lies along that direction.
This is why a filter is sometimes described as a template and the operation as
template matching. The description is right for a single filter and becomes
misleading in deeper layers, where the thing being matched is not in the picture
at all but in the output of earlier filters. The arithmetic is identical; only the
input has changed.
for r in range(H - kh + 1): # every valid row position
for c in range(W - kw + 1): # every valid column position
total = 0.0
for u in range(kh): # the window, same weights every time
for v in range(kw):
total += w[u][v] * x[r + u][c + v]
y[r][c] = totalThe two properties worth naming
Two facts about this operation get used repeatedly later, and both are visible in
the formula. First, it is linear: double the input and the output doubles, add
two inputs and the outputs add. That means the whole thing is a matrix
multiplication in disguise, a claim the final lesson makes explicit.
Second, it commutes with shifting. Move the input two places to the right and the
output moves two places to the right, with every value unchanged. The usual word
for this is equivariance, and it is the precise sense in which a convolution
treats all positions alike. A connected layer has no such property: shifting its
input changes every output, because every input position has its own weights.
| step | window position | values under the window | response |
|---|---|---|---|
| 1 | 0 | 10, 10, 10 | 0 |
| 2 | 1 | 10, 10, 90 | 80 |
| 3 | 2 | 10, 90, 90 | 80 |
| 4 | 3 | 90, 90, 90 | 0 |
| 5 | 4 | 90, 90, 20 | -70 |
What has and has not been decided
So far there is one window, one set of weights, and one map coming out. Three
things have been quietly left open, and each is a lesson of its own. The window
reads values that do not exist when it hangs over the edge of the picture, and
something has to be done about that. It has been slid one place at a time, which
is a choice rather than a necessity. And a single filter detects a single
pattern, where any real use needs dozens at once.
The important thing to carry forward is smaller than any of those. The operation
is a dot product repeated at every position with the same weights, and its output
is a map in the same coordinates as its input. Everything else in this course is
bookkeeping around those two sentences.
Recap
- A convolution output at one position is nothing but a dot product between a small window of weights and the equally small patch of input sitting under it, so the whole operation is built from the most ordinary arithmetic there is.
- Because the same weights are used at every position, the output is not a summary but a map: it has positions of its own, and the value at each one says how strongly the pattern was present at the matching place in the input.
- The operation is linear and it commutes with shifting the input, which together are the two properties that every later lesson in this course leans on.
This is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents