The Browser Starts Building the Page Before It Has Finished Reading It, and One Tag in the Wrong Place Stops the Whole Production Line
Markup becomes a tree while the bytes are still arriving. This lesson follows that conversion, then shows what happens to it the moment the parser meets a script.
Characters, tokens, nodes
A document arrives as a stream of characters with no structure whatsoever. What
the page needs is a tree, because everything downstream, styling, layout and
painting, is defined in terms of parents and children.
The conversion happens in two stages, and separating them is what makes the
whole thing tractable.
A tokeniser reads characters one at a time and emits tokens. A start tag. An
attribute on that tag. A run of text. An end tag. It is a state machine with
no memory of structure, and it does not know or care whether the document makes
sense.
A tree builder consumes those tokens and maintains a stack of currently open
elements. A start tag pushes a node onto the stack and attaches it as a child
of whatever is beneath. Text attaches to whatever is on top. An end tag pops.
That stack is where nesting actually lives.
bytes received so far
<html><head><link rel=stylesheet href=/site.css>
</head><body><h1>Prices</h1><p>From
... connection still open, more coming
tree built so far
html
head
link
body
h1
text: Prices
p
text: From
open element stack: html, body, pBuilding it while it arrives
The important consequence is that none of this waits. A browser receiving the
first kilobyte of a document can tokenise it, build that part of the tree,
resolve styles for it, and in many cases put it on the screen, all while the
rest of the document is still in flight.
| step | bytes in | nodes in the tree | what the browser did with them | on screen | what happened |
|---|---|---|---|---|---|
| 1 | 1400 | 9 | found the stylesheet and started fetchin | nothing | The first packet usually carries the head. Finding the stylesheet reference early is the single most valuable thing in it, because that fetch gates the first paint. |
| 2 | 4200 | 31 | styled the header and worked out its pos | nothing yet, waiting on the stylesheet | The tree is growing and layout can be computed, but painting waits because applying styles later would mean painting the wrong thing first. |
| 3 | 9800 | 74 | stylesheet arrived, painted the header a | header and text | The first frame. Note that less than half the document has arrived, and the user is already reading. |
| 4 | 21000 | 160 | finished the document | the whole page | The last byte changes less than people expect. Most of the perceived load finished several packets earlier. |
This is also why a document that streams beats a document that is assembled
first and sent in one piece. The browser can only work on what it has, so
sending the head immediately and the body as it becomes available gives the
browser a head start measured in whole round trips.
Why a script stops everything
Then the parser reaches a script tag, and the production line stops.
The cost is easy to underestimate because it is not the running of the script
that hurts. It is the fetch. A script tag in the head of a document, with no
attribute, costs a full round trip during which nothing is parsed, nothing is
styled and nothing can be painted.
| blocks parsing while fet | fetched alongside parsin | runs in document order | runs whenever it arrives | |
|---|---|---|---|---|
| plain script tag | 1 | 0 | 1 | 0 |
| marked defer | 0 | 1 | 1 | 0 |
| marked async | 0 | 1 | 0 | 1 |
Reading ahead for resources
One refinement saves the blocking case from being as bad as it sounds. While
the parser is stopped waiting on a script, a second much simpler scanner runs
ahead through the bytes already received, looking only for things worth
fetching: stylesheets, scripts, images.
It does not build a tree and it does not need to be correct about structure. It
only needs to spot references and start the fetches early, so that when the
parser resumes the resources are already on their way.
Two things follow for anybody writing documents.
References the scanner can see get found early. A stylesheet or image named in
the markup is spotted immediately. One that a script decides to fetch after it
runs cannot be, because the scanner does not execute anything, so that resource
starts a round trip or two later than it needed to.
And the scanner is a mitigation, not a cure. The parser is still stopped, the
tree is still not growing, and nothing can be painted. The correct action is
still to not block the parser in the first place.
So the document has become a tree, incrementally, with one well-understood way
of stalling it. The tree on its own places nothing on the screen, because every
node still needs to know what it looks like. That is the next lesson.
Recap
- Parsing is incremental: the tree grows from the bytes already received, so the top of a document can be worked on while the bottom is still in flight.
- A plain script tag stops parsing entirely until the script has been fetched, parsed and run, because the script might change the document being parsed.
- The two attributes that change this are the single highest-value change available in most documents, and which one to use follows from whether order matters.
This is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents