ContentsThe library

Counting Things as They Arrive

The Spike at Nine Was Six Hours of Traffic Arriving at Once

Last timeData With No End

An event has a time it happened and a time it reached you. Each clock gives a wrong answer to the questions the other one owns, and both wrongs look plausible.

The two times on every event

Every record arriving at a streaming system has two times attached to it,

whether or not both were written down.

Event time is when the thing happened: the button was pressed, the sensor took

its reading, the payment cleared. It is set by whatever observed the event,

which is a machine you do not control, running a clock you did not set.

Processing time is when the record reached you. It is read from your own clock

at the moment of arrival, and it is reliable in a way event time is not,

because you own it.

For most records the two are a second or two apart. The gap, usually called

skew, is what the rest of this lesson is about.

FIG 1Skew, which is the only thing worth measuring here
the skew on one record
processing time, when it arrived
event time, when it happened
Trivial arithmetic carrying the most useful column in a streaming pipeline. Store it per record and you can answer how late things are, which sources are worst, how the tail behaves during an incident, and how long a window has to stay open. Almost nobody stores it, and the information cannot be reconstructed later because one of the two times is usually discarded on arrival.
FIG 2Where skew comes from, in a consumer application
Only the middle slice is under your control, and it is the one most teams assume is the whole picture. The largest slice is a property of the world, which means it cannot be engineered away and has to be designed around. The fourth slice is the one that produces events in the future, which is a separate problem taken up at the end of this lesson.

What processing time gets wrong

Group by processing time and the question you are answering is when did I see

these, which is almost never the question anybody asked.

FIG 3A six-hour stall, as each clock reports it
stephourevents really madeevents arrivingby event timeby processing timewhat happened
103:0040000400004000040000Normal operation. Both clocks agree, which is the condition people mistake for the clocks being the same thing.
204:00 to 09:00200000040000 an hourzero an hourThe consumer is stalled. Events are happening and none are arriving. Event time has not noticed; processing time shows the outage exactly.
309:05300020300040000 an hour, filled ina spike of 203000Recovery. Six hours of backlog arrives in five minutes, carrying its original event times.
410:0040000400004000040000Back to normal. By event time the day is flat and correct. By processing time there is a gap followed by a spike, neither of which corresponds to anything a customer did.
4 steps
Read the last two columns as two different graphs of the same day. One shows a business that behaved identically all day, which is true. The other shows a collapse followed by a record-breaking five minutes, which is false as a statement about customers and exactly right as a statement about the pipeline. Neither graph is wrong; they are answers to different questions.

The spike is the part that does damage, because it is indistinguishable from

real traffic. An alert fires on an unusual volume. A capacity decision is made

from a peak that never happened. A fraud rule triggers on a burst of activity

from one region. And any rate computed over that minute is nonsense.

This is also the mechanism behind a subtler failure. A sale made at 23:58 and

received at 00:04 is counted tomorrow under processing time, so yesterday is

short and today is long. Both figures look reasonable and neither is right,

and the size of the error changes with the health of the pipeline, which makes

day-to-day comparison meaningless.

The lesson stops here

3 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01Sort This. There Is No Last Row.
  2. 02The Spike at Nine Was Six Hours of Traffic Arriving at Onceyou are here
  3. 03Three Shapes, and the Question Tells You Which Oneopening only
  4. 04Nothing Can Tell You That Nothing Else Is Comingopening only
  5. 05Somebody Has Already Seen the Number You Are About to Changeopening only
  6. 06An Average Needs Two Numbers, a Median Needs All of Themopening only
  7. 07Twelve Kilobytes Will Count a Billion Different Thingsopening only
  8. 08The Windows Came Back, and So Did Nine Minutes of Totalsopening only

Read alongside