ContentsThe library

What the Planner Decides

It Has Never Looked at Your Data

Last timeThree Ways to Join

Every decision in a plan rests on how many rows each step will produce. The planner works that out from a few hundred numbers gathered hours ago, and here is exactly how.

Everything in the previous lesson came down to sizes. A loop was chosen because

one side was small; a hash was sized for the input it expected. This lesson is

about where those numbers come from, and the answer is that they are guesses

made from a very small summary.

What it actually knows

The planner does not read the table. Reading the table to find out how long

reading the table would take is obviously absurd, so instead a background job

samples the table occasionally and records a summary.

Per table: the number of rows and the average row width. Per column: how many

distinct values there are, how often the column is null, a list of the most

common values with their frequencies, and a histogram dividing the rest of the

range into buckets holding roughly equal numbers of rows.

That is a few hundred numbers per column, standing in for a table of billions

of rows, and gathered from a sample of perhaps thirty thousand of them.

FIG 1A histogram of order amounts, with what each bucket implies
lower boundupper boundrows in the bucketrows at or below the upp
bucket 10129400094000
bucket 2122994000188000
bucket 3295894000282000
bucket 45814094000376000
bucket 5140980094000470000
Five buckets of a real histogram, each holding the same number of rows, which is why the bounds are so uneven. The marked bound is the point: the top bucket spans from 140 to 9800 and holds exactly as many rows as the bucket spanning 0 to 12. Equal height is chosen deliberately, because it puts the resolution where the data is, and it is also the source of the largest errors, since any estimate inside that top bucket assumes an even spread across a range where there plainly is not one.

Estimating one filter

With that summary, a filter is estimated in one of three ways.

If the filter asks for equality against a value in the most common list, the

frequency is read off directly and the estimate is excellent. If it asks for

equality against anything else, the planner assumes the remaining values are

evenly spread and estimates one over the number of distinct values.

The lesson stops here

5 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01You Asked For Rows, Not For a Method
  2. 02One of Them Is Quadratic and Sometimes It Is the Right Answeropening only
  3. 03It Has Never Looked at Your Datayou are here
  4. 04A Factor of Ten at the Bottom Is a Factor of a Thousand at the Topopening only
  5. 05The City Already Told You the Postcodeopening only
  6. 06Six Tables Have Seven Hundred Twenty Orders and They Are Not Alikeopening only
  7. 07Find the Lowest Line Where the Two Numbers Stop Agreeingopening only
  8. 08Four Remedies and the One Cause Each of Them Repairsopening only

Read alongside