ContentsThe library

Matrices as Maps

Not a Grid, a Verb

A matrix is a machine that takes a vector and returns a different one, and the only thing that makes it special is that it respects addition and scaling.

Here is the sentence this course is built on. A matrix is a machine that takes a

vector in and gives a vector out.

That is all. The rectangle of numbers is a record of what the machine does, in

the same way that a dictionary definition is a record of what a word means. You

would not try to understand a word by memorising the arrangement of letters in

its definition, and there is no more reason to try to understand a matrix by

staring at its entries.

FIG 1What the object actually is
Notice that the two routes to the bottom right apply the same pair of maps in opposite orders. In general they do not arrive at the same place, which is the whole reason matrix multiplication is written so carefully, and is the subject of the third lesson.

The two rules

Not every machine that eats vectors and produces vectors is a matrix. Only the

linear ones are, and linear means exactly two things.

FIG 2The entire definition
the map, whatever it does
mapping a sum is the same as mapping the pieces and adding, which is additivity
mapping a scaled vector is the same as scaling the mapped vector, which is homogeneity
exact equality in both lines, for every vector and every number, with no exceptions
Two lines, and the rest of the subject is their consequences. Setting the scale factor to zero in the second line forces the zero vector to map to the zero vector, so the origin can never move under a linear map. That single corollary disposes of a large class of candidate maps.

Checking a map against those two rules is a mechanical exercise, and it is worth

doing a few times because the results are not always the ones people expect.

FIG 3Which of these are linear
additivescaleslinear
swap the two coordinates111
double the first coordin111
square the first coordin010
add three to the first c100
set the second coordinat111
The marked cell is the trap. Adding a constant passes the scaling test for no value of the scale factor except one, and it fails additivity because the constant gets added twice on one side and once on the other. It also moves the origin, which already settles it. Squaring fails scaling, since doubling the input quadruples the output. The last row is linear despite destroying information, which is a point the fourth lesson develops.

The offset case deserves a sentence of its own, because of a clash of

vocabulary. The familiar line with a slope and an intercept is called a linear

equation in school and is not a linear map here. Only the lines through the

origin are. When an intercept is needed, the usual trick is to add a coordinate

that is always one, which turns the offset into part of a genuinely linear map in

a larger space, and that trick is why the bias term in a neural network layer is

sometimes described as a column of the matrix.

One more consequence is worth drawing out, because it is the reason the two

rules are stated together rather than as one. Additivity on its own allows some

badly behaved maps, and scaling on its own allows maps that treat the axes

inconsistently. Taken together they force the map to respect every combination of

vectors at once, which is precisely the structure the next lesson exploits to

read a matrix off a short list of destinations.

What it looks like

Geometrically the two rules are a strong constraint, and the picture is worth

carrying.

FIG 4What survives a linear map and what does not
stepfeature of spaceafter a linear mapwhy
1straight linesstill straighta line is a point plus multiples of a di
2parallel linesstill parallelthey share a direction, and that directi
3equal spacing along a linestill equalscaling is preserved, so equal steps map
4the originstill at the originforced by the scaling rule with a factor
5lengths and anglesgenerally destroyednothing in the two rules protects them,
5 steps
The last row is the one people forget. A linear map is free to stretch one direction and squash another, so distances and angles mean something different after it. The grid of space stays a grid, but the cells may be larger, smaller, slanted, or flattened to nothing.

Why bother restricting to these

The restriction looks severe, and it is. Most functions are not linear. So why

does an entire branch of mathematics and most of practical computing organise

itself around the ones that are?

FIG 5What the restriction buys
the output for any input at all, reconstructed from the axis images alone
where the map sends the first axis direction, which is a finite amount of information
a coordinate of the input, saying how much of that axis image to use
where the map sends the last axis, after which there is nothing left to record
This line is the payoff. Write any vector as a combination of the axes, apply the two rules, and the map pulls apart into what it does to each axis. So knowing a handful of outputs determines the map everywhere. Nothing remotely like this is true of a general function, where knowing a million outputs tells you almost nothing about the next one.

That is the bargain. Give up bending and shifting, and in exchange the whole map

collapses into a finite table that a computer can store and multiply quickly.

Enormously complicated systems are built by stacking these simple maps and

inserting a small amount of bending between them, which is a fair description of

a neural network.

The next lesson reads that table. If the map is determined by where it sends the

axes, then the columns of the matrix are exactly those destinations, and reading

a matrix becomes a matter of looking at its columns.

Recap

  • A matrix is a function from vectors to vectors, and the grid of numbers is only a record of what that function does.
  • Linear means two things: the map of a sum is the sum of the maps, and the map of a scaled vector is the scaled map. Everything else follows from those two.
  • Linearity is a severe restriction, ruling out bending and shifting, and it is exactly what makes a map describable by a finite grid of numbers.

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

NextReading a Matrix Off Its Columns →

The rest of this course

  1. 01Not a Grid, a Verbyou are here
  2. 02Each Column Is a Destinationopening only
  3. 03Why the Multiplication Rule Looks Like Thatopening only
  4. 04The Shadow and the Thingopening only
  5. 05Undo, If You Canopening only
  6. 06The Directions That Only Stretchopening only
  7. 07Every Map Is a Rotation, a Stretch and a Rotationopening only
  8. 08When You Cannot Hit the Targetopening only

Read alongside