Six Tables Have Seven Hundred Twenty Orders and They Are Not Alike
Last timeColumns That Are Not Independent
Which tables are joined first decides how many rows every later step carries. The orders cannot all be tried, so they are built up in pieces and reused.
Estimation is behind us. This lesson and the next are about the choices
themselves, and this is the one that matters most: which tables are joined
first.
Why the order decides the cost
Every join method in the second lesson costs something proportional to the size
of what goes into it. A nested loop is linear in the outer input. A hash join
pays to build from one side and probe with the other. A merge join sweeps both.
None of them is cheap on a large input and all of them are cheap on a small one.
Now notice that the input to the second join is not a table. It is whatever came
out of the first. The input to the third is whatever came out of the second. So
the order fixes the size of every input except the very first, and the sizes of
those intermediate results, not the sizes of the tables, are what the work is
proportional to.
Putting it plainly: a filter that removes ninety nine per cent of a table is
worth almost nothing if it is applied after that table has been joined to three
others, and worth almost everything if it is applied before.
The lesson stops here
6 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents