Complaints That Are Really About This Layer
Last timeThe Same Word Twice
Several familiar model failures are not failures of reasoning at all but consequences of how the text was cut, and they can be told apart with a simple test.
A model multiplies two long numbers and gets it wrong. A model is asked how many
times a letter appears in a word and miscounts. A model handles a request in one
language well and the same request in another badly. These look like three
different shortcomings and they share one cause.
Take the arithmetic first. Column arithmetic requires lining up digits by place
value. The model does not receive digits. It receives whatever digit groups the
list happens to contain, and those groups depend on which digit sequences were
frequent in the corpus, so the same quantity can be split in different places
depending on its digits. Getting place values to line up when the two numbers were
chopped differently is extra work that has nothing to do with arithmetic.
| step | stage | what the model has | what the task needs |
|---|---|---|---|
| 1 | the request as text | a string of digits | two numbers |
| 2 | after cutting | a few digit groups of uneven length | digits, individually |
| 3 | identifying place value | nothing marking which group is the thous | an explicit column position per digit |
| 4 | aligning the two numbers | groups that may be split at different po | columns that correspond between the two |
| 5 | carrying between columns | carries crossing group boundaries | a carry between adjacent digits |
Letters inside a piece
The spelling complaint is a cleaner version of the same thing. If a word arrives as
a single piece, its letters are not in the input. There is nothing to count. The
model can often answer anyway, because the spelling of common words is something it
has encountered in text, but that is recall rather than inspection, and recall
fails on exactly the words you would expect.
This also explains the pattern of successes. Common words are frequently answered
correctly and unusual ones are not, which is the opposite of what a counting
procedure would do, and it is a clue that no counting is happening.
Code and indentation
Code puts weight on characters that prose treats as filler. Leading whitespace
determines the structure of a program in several widely used languages, and a
tokeniser built mostly on prose has entries for runs of spaces reflecting how prose
uses them.
The result is that indentation gets cut into pieces that do not correspond to one
level of nesting, so the model has to infer structure from a representation that
obscures it. Tokenisers intended for code add entries for the common indentation
runs for exactly this reason, which is one of the clearest cases of the list of
pieces being tuned to a domain rather than to language in general.
The lesson stops here
3 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents