ContentsThe library

Designing an Interface to Call

Counting From the Start Versus Remembering the Place

Last timeWhen the Caller Calls Twice

Two ways to hand back a long list. One is easier to build and quietly loses rows while somebody is reading. The difference is worth understanding before you choose.

Why not return all of it

An operation that returns a collection has a decision to make that an

operation returning one thing does not: how much. The tempting answer is all

of it, and it works beautifully in testing, where the collection has nine

rows.

What goes wrong is not only speed. The response size depends on data you do

not control, which means it has no upper bound, which means every number

downstream of it has no upper bound either: the memory your service uses to

build it, the memory the caller uses to parse it, the time on the wire, the

chance of hitting a timeout. And the usual response to a timeout makes it

worse, because the caller retries the same enormous request and you do the

same enormous work again.

So the size is capped, by you, as part of the design. That immediately

creates the real question, which is how the caller asks for the next part.

FIG 1The shape of one page
plaintext
GET /invoices?limit=50&starting_after=inv_8KQ2mz

{
  data: [ fifty invoices, newest first ],
  has_more: true,
  next: /invoices?limit=50&starting_after=inv_4Rb7pq
}

the caller loop
  ask for the first page with no position
  read data
  if has_more is false, stop
  otherwise ask for next, unchanged
  repeat

what the caller never does
  parse the position value
  add to it, or guess the next one
  keep a count of how many rows have gone past
The request takes a limit and a position. The response hands back the next position rather than making the caller build it, which keeps the construction of that value entirely on your side and lets you change it later. The has_more flag exists so an empty final page is not required to end the loop.

Counting from the start

The scheme everybody writes first is a count and a size: skip the first

hundred rows, give me twenty. It is easy to implement, easy to explain, and

it gives the caller something genuinely useful, which is the ability to jump

to any page directly.

It has one assumption buried in it, and the assumption is that the rows

before the reader's position do not change while they are reading. Stated

that way it is obviously false for any collection being written to.

FIG 2One row inserted while somebody reads
stepmomentrows in the collectionthe reader asks forwhat they getwhat happened
1startA B C D E Fskip 0, take 2A BCorrect. The reader has seen two rows and believes the next two start at position 2.
2a write landsZ A B C D E Fnothing yetnothing yetA new row arrives at the front, which is the common case for anything ordered newest first. Every row has moved down one place.
3second requestZ A B C D E Fskip 2, take 2B CPosition 2 now holds B, which the reader already has. So B arrives twice and A is still fine.
4after a deletion insteadA C D E Fskip 2, take 2D EWith B deleted rather than Z inserted, position 2 holds D. The row C is never returned to anybody, and nothing anywhere reports an error.
4 steps
Two failures from one cause. An insertion before the reader duplicates a row, a deletion before them skips one, and the skip is the serious one because the reader has no way to notice it. A job consuming the whole collection quietly processes less than all of it.

The second problem with counting is cost, and it is a different kind of

problem because it is only about speed. To skip a hundred thousand rows a

database generally has to walk a hundred thousand rows. The deeper the reader

goes, the slower each page gets, which is the opposite of what anybody

expects.

The lesson stops here

3 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01Every Awkward Interface You Have Ever Used Was Awkward for the Same Reason, and the Reason Was Decided in the First Half Hour
  2. 02There Are Only Two Questions Worth Asking About Any Operation, and Neither of Them Is What It Is Calledopening only
  3. 03Your Error Message Is Read by a Program First and a Human Second, and Almost Every Interface Gets That Order Backwardsopening only
  4. 04The Second Call Is Not a Mistakeopening only
  5. 05Counting From the Start Versus Remembering the Placeyou are here
  6. 06A Limit Is Only Useful If Somebody Can Obey Itopening only
  7. 07Hand Back a Receipt, Not an Answeropening only
  8. 08Add, Migrate, Then Removeopening only

Read alongside