Derive why gradient descent works, what breaks it, and what every optimiser since is repairing, well enough to read a training curve and know which number to change
The Mathematics of Gradient Descent

A derivation-first walk through training, taking the steepest direction, the largest safe step, momentum, noise and per-coordinate scaling each in turn.
8 lessons, written and corrected before you arrived. Reading them here needs no account. The first reads the whole way through; the others open and then stop. Starting the course gives you your own copy, where every idea has problems standing under it and you can ask about any sentence.