ContentsThe library

The Mathematics of Attention

Where the Square Root of dkd_k Comes From

Last timeQuery, Key and Value

The division by the square root of the key dimension is not a tuning constant. It falls out of the variance of a dot product between independent vectors.

Of all the pieces of the attention formula, the dk\sqrt{d_k} in the denominator

looks most like something someone tried until training stopped diverging. It is

not. It is the standard deviation of the quantity above it, and the derivation

takes about five lines.

The setup

Assume, for the length of this argument, that the components of a query vector

qq and a key vector kk are independent random variables with mean zero and

variance one. This is roughly what a freshly initialised network gives you, and

roughly what normalisation maintains during training.

A single attention score is their dot product:

q⋅k=∑i=1dkqikiq \cdot k = \sum_{i=1}^{d_k} q_i k_i

The question is how large this number typically is, and specifically how its

size depends on dkd_k.

The lesson stops here

11 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01Why Attention Needs Three Matrices, Not One
  2. 02Where the Square Root of dkd_k Comes Fromyou are here
  3. 03Why Softmax and Not Some Other Normalisationopening only
  4. 04One Weighted Average Is Not Enoughopening only
  5. 05Attention Cannot See Order, So Order Has to Be Addedopening only
  6. 06The Skip Connection Is the Main Roadopening only
  7. 07Why Transformers Normalise Across Features and Not Across the Batchopening only
  8. 08The Half of the Transformer That Does Not Move Informationopening only

Read alongside