Derive the loss almost every classifier and language model is trained on, from the probability of the data through to what its value does and does not promise, well enough to read a loss of 1.8 and say what the model believes
The Mathematics of Cross Entropy

A derivation-first account of the training loss itself, from likelihood and the logarithm through softmax, the gradient, the entropy floor, perplexity and calibration.
8 lessons, written and corrected before you arrived. Reading them here needs no account. The first reads the whole way through; the others open and then stop. Starting the course gives you your own copy, where every idea has problems standing under it and you can ask about any sentence.