Follow a derivative all the way from a scalar loss back to every weight in a deep network, with the rule for each kind of operation, the memory the method costs, and the reasons the signal can arrive too small or too large
Backpropagation in Full

One number at the end, millions of weights at the start, and a single sweep that connects them. This course derives backpropagation from the chain rule, works out the rule for each layer, and shows what it costs in memory.
8 lessons, written and corrected before you arrived. Reading them here needs no account. The first reads the whole way through; the others open and then stop. Starting the course gives you your own copy, where every idea has problems standing under it and you can ask about any sentence.