Designing an Interface to Call
Counting From the Start Versus Remembering the Place
Last timeWhen the Caller Calls Twice
Two ways to hand back a long list. One is easier to build and quietly loses rows while somebody is reading. The difference is worth understanding before you choose.
Why not return all of it
An operation that returns a collection has a decision to make that an
operation returning one thing does not: how much. The tempting answer is all
of it, and it works beautifully in testing, where the collection has nine
rows.
What goes wrong is not only speed. The response size depends on data you do
not control, which means it has no upper bound, which means every number
downstream of it has no upper bound either: the memory your service uses to
build it, the memory the caller uses to parse it, the time on the wire, the
chance of hitting a timeout. And the usual response to a timeout makes it
worse, because the caller retries the same enormous request and you do the
same enormous work again.
So the size is capped, by you, as part of the design. That immediately
creates the real question, which is how the caller asks for the next part.
GET /invoices?limit=50&starting_after=inv_8KQ2mz
{
data: [ fifty invoices, newest first ],
has_more: true,
next: /invoices?limit=50&starting_after=inv_4Rb7pq
}
the caller loop
ask for the first page with no position
read data
if has_more is false, stop
otherwise ask for next, unchanged
repeat
what the caller never does
parse the position value
add to it, or guess the next one
keep a count of how many rows have gone pastCounting from the start
The scheme everybody writes first is a count and a size: skip the first
hundred rows, give me twenty. It is easy to implement, easy to explain, and
it gives the caller something genuinely useful, which is the ability to jump
to any page directly.
It has one assumption buried in it, and the assumption is that the rows
before the reader's position do not change while they are reading. Stated
that way it is obviously false for any collection being written to.
| step | moment | rows in the collection | the reader asks for | what they get | what happened |
|---|---|---|---|---|---|
| 1 | start | A B C D E F | skip 0, take 2 | A B | Correct. The reader has seen two rows and believes the next two start at position 2. |
| 2 | a write lands | Z A B C D E F | nothing yet | nothing yet | A new row arrives at the front, which is the common case for anything ordered newest first. Every row has moved down one place. |
| 3 | second request | Z A B C D E F | skip 2, take 2 | B C | Position 2 now holds B, which the reader already has. So B arrives twice and A is still fine. |
| 4 | after a deletion instead | A C D E F | skip 2, take 2 | D E | With B deleted rather than Z inserted, position 2 holds D. The row C is never returned to anybody, and nothing anywhere reports an error. |
The second problem with counting is cost, and it is a different kind of
problem because it is only about speed. To skip a hundred thousand rows a
database generally has to walk a hundred thousand rows. The deeper the reader
goes, the slower each page gets, which is the opposite of what anybody
expects.
The lesson stops here
3 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents