Not a Grid, a Verb
A matrix is a machine that takes a vector and returns a different one, and the only thing that makes it special is that it respects addition and scaling.
Here is the sentence this course is built on. A matrix is a machine that takes a
vector in and gives a vector out.
That is all. The rectangle of numbers is a record of what the machine does, in
the same way that a dictionary definition is a record of what a word means. You
would not try to understand a word by memorising the arrangement of letters in
its definition, and there is no more reason to try to understand a matrix by
staring at its entries.
The two rules
Not every machine that eats vectors and produces vectors is a matrix. Only the
linear ones are, and linear means exactly two things.
- the map, whatever it does
- mapping a sum is the same as mapping the pieces and adding, which is additivity
- mapping a scaled vector is the same as scaling the mapped vector, which is homogeneity
- exact equality in both lines, for every vector and every number, with no exceptions
Checking a map against those two rules is a mechanical exercise, and it is worth
doing a few times because the results are not always the ones people expect.
| additive | scales | linear | |
|---|---|---|---|
| swap the two coordinates | 1 | 1 | 1 |
| double the first coordin | 1 | 1 | 1 |
| square the first coordin | 0 | 1 | 0 |
| add three to the first c | 1 | 0 | 0 |
| set the second coordinat | 1 | 1 | 1 |
The offset case deserves a sentence of its own, because of a clash of
vocabulary. The familiar line with a slope and an intercept is called a linear
equation in school and is not a linear map here. Only the lines through the
origin are. When an intercept is needed, the usual trick is to add a coordinate
that is always one, which turns the offset into part of a genuinely linear map in
a larger space, and that trick is why the bias term in a neural network layer is
sometimes described as a column of the matrix.
One more consequence is worth drawing out, because it is the reason the two
rules are stated together rather than as one. Additivity on its own allows some
badly behaved maps, and scaling on its own allows maps that treat the axes
inconsistently. Taken together they force the map to respect every combination of
vectors at once, which is precisely the structure the next lesson exploits to
read a matrix off a short list of destinations.
What it looks like
Geometrically the two rules are a strong constraint, and the picture is worth
carrying.
| step | feature of space | after a linear map | why |
|---|---|---|---|
| 1 | straight lines | still straight | a line is a point plus multiples of a di |
| 2 | parallel lines | still parallel | they share a direction, and that directi |
| 3 | equal spacing along a line | still equal | scaling is preserved, so equal steps map |
| 4 | the origin | still at the origin | forced by the scaling rule with a factor |
| 5 | lengths and angles | generally destroyed | nothing in the two rules protects them, |
Why bother restricting to these
The restriction looks severe, and it is. Most functions are not linear. So why
does an entire branch of mathematics and most of practical computing organise
itself around the ones that are?
- the output for any input at all, reconstructed from the axis images alone
- where the map sends the first axis direction, which is a finite amount of information
- a coordinate of the input, saying how much of that axis image to use
- where the map sends the last axis, after which there is nothing left to record
That is the bargain. Give up bending and shifting, and in exchange the whole map
collapses into a finite table that a computer can store and multiply quickly.
Enormously complicated systems are built by stacking these simple maps and
inserting a small amount of bending between them, which is a fair description of
a neural network.
The next lesson reads that table. If the map is determined by where it sends the
axes, then the columns of the matrix are exactly those destinations, and reading
a matrix becomes a matter of looking at its columns.
Recap
- A matrix is a function from vectors to vectors, and the grid of numbers is only a record of what that function does.
- Linear means two things: the map of a sum is the sum of the maps, and the map of a scaled vector is the scaled map. Everything else follows from those two.
- Linearity is a severe restriction, ruling out bending and shifting, and it is exactly what makes a map describable by a finite grid of numbers.
This is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents