ContentsThe library

How Text Becomes Numbers

Complaints That Are Really About This Layer

Last timeThe Same Word Twice

Several familiar model failures are not failures of reasoning at all but consequences of how the text was cut, and they can be told apart with a simple test.

A model multiplies two long numbers and gets it wrong. A model is asked how many

times a letter appears in a word and miscounts. A model handles a request in one

language well and the same request in another badly. These look like three

different shortcomings and they share one cause.

Take the arithmetic first. Column arithmetic requires lining up digits by place

value. The model does not receive digits. It receives whatever digit groups the

list happens to contain, and those groups depend on which digit sequences were

frequent in the corpus, so the same quantity can be split in different places

depending on its digits. Getting place values to line up when the two numbers were

chopped differently is extra work that has nothing to do with arithmetic.

FIG 1Why a multiplication is harder than it looks
stepstagewhat the model haswhat the task needs
1the request as texta string of digitstwo numbers
2after cuttinga few digit groups of uneven lengthdigits, individually
3identifying place valuenothing marking which group is the thousan explicit column position per digit
4aligning the two numbersgroups that may be split at different pocolumns that correspond between the two
5carrying between columnscarries crossing group boundariesa carry between adjacent digits
5 steps
Each row names work the model has to do before any arithmetic starts, and none of it would be necessary if digits arrived one at a time. This is why models that deliberately cut numbers into single digits are measurably better at arithmetic, and why putting spaces between digits in a prompt sometimes rescues a calculation.

Letters inside a piece

The spelling complaint is a cleaner version of the same thing. If a word arrives as

a single piece, its letters are not in the input. There is nothing to count. The

model can often answer anyway, because the spelling of common words is something it

has encountered in text, but that is recall rather than inspection, and recall

fails on exactly the words you would expect.

This also explains the pattern of successes. Common words are frequently answered

correctly and unusual ones are not, which is the opposite of what a counting

procedure would do, and it is a clue that no counting is happening.

FIG 2Diagnosing which layer failed
This test is worth making a habit because the two kinds of failure call for opposite responses. A cutting problem is fixed by changing how you write the input, which is free. A capability problem is not fixed that way, and treating one as the other wastes effort in both directions.

Code and indentation

Code puts weight on characters that prose treats as filler. Leading whitespace

determines the structure of a program in several widely used languages, and a

tokeniser built mostly on prose has entries for runs of spaces reflecting how prose

uses them.

The result is that indentation gets cut into pieces that do not correspond to one

level of nesting, so the model has to infer structure from a representation that

obscures it. Tokenisers intended for code add entries for the common indentation

runs for exactly this reason, which is one of the clearest cases of the list of

pieces being tuned to a domain rather than to language in general.

The lesson stops here

3 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01The Layer Nobody Looks At
  2. 02Three Places You Could Cutopening only
  3. 03Nobody Wrote This Listopening only
  4. 04The Number Somebody Had To Chooseopening only
  5. 05One Row, Looked Upopening only
  6. 06What Ended Up In The Tableopening only
  7. 07The Row Is Only The Starting Pointopening only
  8. 08Complaints That Are Really About This Layeryou are here

Read alongside