QMIND Research · Deck 2 of 7
Neural networks from scratch
A handwritten digit, 13,002 numbers, and the loop that tunes them.
The storyline follows 3Blue1Brown's Neural Networks series, chapters 1–4, by Grant Sanderson. Words and visuals here are ours; the videos are linked at the end.
~1 min
This deck builds one tiny network end to end: what it computes, how we score it, how it learns. Every later deck (transformers, our U-Net) runs on the same machinery.
Credit up front: the arc is Grant Sanderson's (3Blue1Brown) four-chapter series. If you only watch one thing after today, watch those.
Everything animated in this deck comes from a real 784-16-16-10 network trained in numpy for this talk: real weights, real digits, real gradients.
Intuition: a neural network is a big adjustable function; learning means turning its knobs to make a "wrongness" score go down.
If asked: "Why such a small network?" Small enough that we can look at every neuron. The ideas are exactly what scales up to billions of parameters.
The puzzle
You read a sloppy 3 instantly. Could you code it?
To a computer: three grids of 784 numbers that barely agree.
So we won't write rules. We'll tune a function on examples.
~1.5 min
Three real handwritten 3s from MNIST. Your visual cortex says "3" before you can even think about it.
Stack them: each tinted differently. Where they overlap is tiny; the count at the bottom is the number of pixels lit in all three.
So "is this a 3?" is not a pixel lookup. Try writing the if-statements: thickness, slant, size, position all vary.
The move of machine learning: write a flexible function with many knobs, then let examples set the knobs.
Intuition: the same concept can be wildly different at the pixel level; the network has to learn what stays the same.
Analogy: like writing a spam filter by hand vs. letting it learn from labelled email.
If asked: "What's MNIST?" 70,000 28×28 grayscale handwritten digits (60k train, 10k test), the classic benchmark since the 1990s. Our run used 10,000 of them bundled in an npm package.
Roadmap
Four ideas, one tiny network
Part 1
Structure Neurons, layers, weights, biases. The network is a function.
Part 2
Gradient descent Score it with a cost; walk downhill. What did it learn?
Part 3
Backpropagation Which knobs to turn, and how much, for one example.
Part 4
The calculus The chain rule that makes Part 3 exact.
Same order as 3Blue1Brown chapters 1–4. Links on the last slides.
~1 min
Part 1: what the network is : a function from 784 pixel values to 10 scores.
Part 2: how to measure "wrong" and make it less wrong. Then we open up the trained network and get a surprise.
Part 3: backprop as intuition: for one example, who should change and by how much.
Part 4: the same thing as calculus. It's the chain rule, nothing more exotic.
Tell people: you'll be presenting one of these topics tomorrow, so note which part you'd like.
Intuition: structure → objective → optimizer → how to compute gradients. That's every deep learning system.
PART 1
What is a neural network?
Neurons that hold numbers, layers that feed each other, and 13,002 knobs.
Structure
A neuron is just a thing that holds a number
Its activation : how lit up it is.
0 = dark, 1 = fully on.
Everything else is rules for computing these numbers.
~1 min
Strip away the mystique: a neuron is a container for one number.
That number is called its activation. Here it's 0.82: fairly lit up.
In this network every activation lives between 0 and 1. Brighter = bigger.
The rest of Part 1 is: where do these numbers come from, and how does one layer's numbers decide the next?
Intuition: a neural network is a pile of numbers computed from other numbers.
If asked: "Is it always between 0 and 1?" Only with sigmoid. With ReLU (most modern nets) activations are any non-negative number; we'll see that in a few slides.
Structure · input layer
A 28×28 image becomes 784 input neurons
28 × 28 = 784 pixels, each a brightness from 0 to 1
Unroll them into one column
One neuron per pixel
~1.5 min
A real MNIST digit. Each square is one pixel; the image is just a 28×28 grid of numbers.
Zoom: these are the actual brightness values, 0.00 for black up to 1.00 for white.
Unroll the grid row by row into one long column of 784 numbers. Watch the rows fly in order.
Each entry of the column is an input neuron. Its activation is just the pixel's brightness. No computation yet.
Intuition: the network never sees "an image", only a list of 784 numbers.
If asked: "Doesn't flattening throw away which pixels are neighbours?" Yes! This network has to rediscover spatial structure from data. Convolutional nets (and our U-Net) keep the 2D layout built in, which is a big reason they work better on images.
Structure · layers
Two hidden layers of 16, then ten outputs
10 outputs , one per digit: how strongly it thinks "this is a 7".
2 hidden layers of 16. An arbitrary choice: small enough to watch.
Every neuron feeds every neuron in the next layer.
The answer = the brightest output.
~1.5 min
The column of 784 inputs from the previous slide is our first layer.
Last layer: 10 neurons, one per digit. Their activations are the network's votes.
In between, two "hidden" layers of 16. Why 2 and 16? No deep reason; it's a hyperparameter. 3Blue1Brown chose it so the picture stays readable, and we kept it.
Fully connected: each neuron listens to all neurons in the layer before. That's 784×16 lines just into the first hidden layer.
To classify: run the image through and pick the most active output.
Intuition: information flows left to right; each layer is a new description of the image.
If asked: "Why call them hidden?" We only specify inputs (pixels) and desired outputs (labels); what the middle layers do is up to training.
Structure · forward pass
Each layer's activations determine the next
Real network, real weights. Brighter line = bigger weight times activation. Blue pushes the next neuron up, rose pushes it down.
~1.5 min
This is our trained network, the actual weights, running live in the browser.
Load the pixels into the input layer.
Every input pushes on every first-layer neuron. Brightness of a line = how much it contributes for this image: blue pushes the neuron up, rose pushes it down.
Same again into the second hidden layer.
And into the outputs. The brightest output is the answer (highlighted in amber).
A different digit, a different ripple. Same weights, completely different pattern of activity.
One more. It's all one fixed function; only the input changed.
Intuition: the weights are frozen; the image decides which paths light up.
If asked: "Are the outputs probabilities?" Not exactly here: sigmoid outputs don't sum to 1. Modern classifiers use softmax so they do; deck 3 uses softmax for next-token probabilities.
Structure · the hope
The hope: layers build edges into parts into digits
Reusable parts: an 8 is two loops, a 6 is a loop and a line.
Spoiler: our trained net does not do this cleanly
~1.5 min
Why layers at all? Here's the story people hope is true.
A 9 is a loop on top of a line.
So maybe one second-layer neuron lights up for "loop at the top", another for "vertical line on the right".
And each part is made of short edges: maybe first-layer neurons detect little edge pieces.
Edges → parts → digits. Parts are reusable across digits, which is what makes layering efficient. This picture is our illustration, not something we measured.
Remember this hope. In Part 2 we'll open the trained network and see that it mostly doesn't work this way.
Intuition: depth lets you compose simple features into complex ones.
If asked: "Do any networks learn this?" Large convolutional nets do learn something close: edge and texture detectors early, parts later. Tiny fully connected nets mostly don't.
Structure · one neuron
One neuron: weigh every pixel, then add
Positive weights where you want ink, negative around it: a detector for a horizontal stroke.
~2 min
Zoom in on a single first-layer neuron. It looks at all 784 pixels.
It has one weight per pixel. Drawn as an image: blue = positive, rose = negative, dark = zero. These weights are hand-designed for illustration.
Multiply each pixel by its weight. Black pixels contribute nothing, whatever their weight.
Add up all 784 products: one number, the weighted sum. Big when the ink lands on the blue region.
Blue band with rose borders = "stroke here, and nothing just above or below it". That's an edge detector. Weights are a pattern the neuron is looking for.
Intuition: a weighted sum is a template match: dot product of the image with the weight pattern.
Analogy: a stencil held over the image: count ink in the holes, subtract ink on the solid border.
If asked: "Is it a dot product?" Exactly: . Large when the two vectors point the same way.
Structure · squish
Squish the sum into (0, 1); the bias sets the bar
Very negative → near 0. Very positive → near 1.
Bias : how big the sum must be before the neuron lights up.
~1.5 min
The weighted sum can be any number: −20, 3.7, 150. But we want an activation between 0 and 1.
The sigmoid (logistic) function: smooth, S-shaped, always between 0 and 1.
Watch the dots: each sum gets squished onto the 0-to-1 axis. Big negatives pile up near 0, big positives near 1, and only the middle region really distinguishes values.
Add a bias before squishing. With the whole curve shifts right.
Now the neuron only lights up when the weighted sum beats 5. The bias is a threshold knob. Each neuron has its own.
Intuition: weights say what pattern the neuron looks for; the bias says how strongly it must be present.
If asked: "Why squish at all?" Without a nonlinearity, stacking layers collapses to one big linear map: . The squish is what makes depth worth anything.
Structure · modern default
Modern networks mostly use ReLU instead
Sigmoid's slope is at most and nearly 0 in the tails, so learning signals fade layer after layer.
ReLU's slope is 1 wherever it's on. Today's nets use it or smooth cousins (GELU, SiLU). We keep sigmoid to match the classic.
~1 min
Sigmoid was the classic choice, loosely inspired by neurons being off or on.
ReLU, the rectified linear unit: zero for negatives, the identity for positives. Dead simple.
The problem with sigmoid: its slope is tiny in the tails (0.05 at x = 3). In Part 4 we'll multiply slopes together across layers; multiply many small numbers and the gradient vanishes.
ReLU's slope is exactly 1 when active, so gradients pass through deep stacks. Transformers use GELU/SwiGLU-style variants; diffusion U-Nets commonly use SiLU. Same idea: mostly-linear, cheap, good gradients.
Intuition: a nonlinearity you can push gradients through is worth more than one that looks biological.
If asked: "Doesn't ReLU have zero slope for negatives?" Yes, so a neuron can "die" if it's always negative. Leaky ReLU and GELU soften this; in practice it's manageable.
Structure · in matrix form
A whole layer is one matrix–vector product
Toy numbers. The real first layer: is 16 × 784. GPUs are built to do exactly this.
~1.5 min
Writing 16 separate weighted sums is messy. Pack all the weights into a matrix: row holds the weights of neuron .
Row 1 times the activation vector: multiply entry by entry, add. That's exactly the weighted sum for neuron 1.
Sweep down the rows: each row produces one neuron's sum. Matrix times vector = all the weighted sums at once.
Add the bias vector, one bias per neuron.
Apply sigmoid to each entry. Out comes the next layer's activation vector. The whole network is this, three times.
Intuition: one layer = matrix multiply, add, squish. Code and hardware are optimised for exactly that.
If asked: "Why does matrix form matter?" Speed: GPUs and TPUs are essentially matrix-multiply machines, and batching many images turns this into a matrix–matrix product. Deck 4 and 7 go deep on this.
Structure · count the knobs
Count the knobs: 13,002 parameters
The whole network is one function with 13,002 knobs .
~1 min
Every line into the first hidden layer is a weight: 784 × 16 = 12,544.
Hidden to hidden: 256.
Hidden to output: 160.
One bias per non-input neuron: 42.
Total 13,002. Note the bars: 96% of the knobs sit in the first layer, because the input is wide.
So: a function with 784 inputs, 10 outputs, and 13,002 parameters . Complicated, but just a function. "Learning" will mean choosing .
Intuition: structure is fixed by us; the 13,002 numbers are what training finds.
If asked: "How does that compare?" Our U-Net will have millions; frontier LLMs have hundreds of billions to trillions. Same kind of knob.
PART 2
How a network learns: gradient descent
Score how wrong it is. Then walk downhill.
~20 s
We have a function with 13,002 knobs set at random. How do we find good settings? Two ingredients: a score, and a way to improve it.
Gradient descent · cost
Score one guess: the cost of a single example
Trained: small cost means confident and correct.
~1.5 min
Start with random weights (our network's actual starting point). Feed in a 3. The output is mush: these are the real activations.
What we wanted: 1 on the "3" output, 0 everywhere else. That's , the one-hot target, in green.
The gap on each output, squared. Squaring makes all errors positive and punishes big misses more.
Add the ten squares: one number, the cost of this example. Big = bad.
Now swap in the trained weights: outputs snap toward the target and the cost collapses.
Intuition: the cost turns "how wrong" into a single number we can try to shrink.
If asked: "Is squared error what people use?" For classification, modern nets use softmax + cross-entropy, which trains faster. Squared error is what 3Blue1Brown uses because it's simplest; we trained ours with it too, to keep the slides honest.
Gradient descent · cost
Average the cost over every training example
Network
784 pixels in → 10 numbers out
Cost
13,002 knobs in → 1 number out
Learning = find knobs that make it small.
~1.5 min
One example isn't enough; we care about all of them. Here are 20 real digits.
Each one's cost under the random network (rose bars, real values).
Average them: the cost . Our real run averaged over 8,500 training images.
Subtle reframe: the network takes pixels and gives 10 numbers. The cost takes the 13,002 knobs (with data fixed) and gives one number: how bad those knobs are.
With the trained knobs, every bar shrinks. Learning is minimising this function of 13,002 variables.
Intuition: training is optimisation; the data is the landscape, the knobs are your position.
If asked: "Doesn't minimising training cost just memorise?" It can. That's why we score on held-out test images (94% here) and why deck 6 is all about held-out splits.
Gradient descent · one knob
Minimise by walking downhill
Slope says which way is down.
Steps shrink as it flattens out.
Different start, different valley. No guarantee of the best one.
~1.5 min
Pretend the cost depended on one knob . We can't solve for the minimum in closed form for a real network, so we feel our way.
Compute the slope where we stand. Positive slope → go left; negative → go right.
Step size times the slope. is the learning rate.
Because the step is proportional to the slope, steps shrink automatically as the ground flattens.
Keep going until it settles in a valley: a local minimum.
Start elsewhere and you can land in a different, lower valley. Gradient descent only promises "downhill", not "the best".
Intuition: you're blindfolded on a hillside; you can only feel the slope under your feet.
If asked: "Are local minima a big problem in practice?" Less than you'd fear in very high dimensions: most critical points are saddles, and many minima are about equally good. Initialisation and learning-rate schedules matter more.
Gradient descent · two knobs
With two knobs: step against the gradient
points uphill ; is the steepest way down .
Small steps: steady progress.
Too big: overshoot and zigzag.
Same idea with 13,002 knobs. We just can't draw it.
~1.5 min
Now the cost depends on two knobs: a landscape seen from above. Contour lines join points of equal cost; the minimum is the green cross.
The gradient is the vector of partial derivatives; it points in the direction of steepest increase. Flip it to go down fastest.
Step, recompute, step. Steady. Notice it crosses the narrow direction fast and crawls along the long valley.
Crank up the learning rate: it overshoots across the valley and zigzags. Too big and it diverges entirely.
Our network: identical idea, but the "position" is a point in 13,002-dimensional space.
Intuition: the gradient is a local compass; the learning rate is your stride length.
If asked: "How do people fix the zigzag?" Momentum (we used it) and adaptive optimisers like Adam average out the back-and-forth.
Gradient descent · the gradient
: 13,002 nudges, and some matter far more
Our real gradient at the start of training. Blue : turn this knob up. Rose : turn it down.
32% are exactly zero : weights on pixels that were blank in every image of this batch.
Sorted by size: a few nudges dwarf the rest.
Top 1% (130 entries) carry 40% of the total size. The biggest is ~320× the median nonzero nudge.
~1.5 min
Every slot is one of the 13,002 parameters, laid out in columns. This is the real negative gradient of our network at initialisation, averaged over a batch of 100 training images.
Colour shows the sign: blue means increasing that knob lowers the cost; rose means decreasing it does. Brightness is on a compressed scale so small values stay visible.
A third are exactly zero: a weight attached to a pixel that's always black has no effect on the output, so no nudge. Mostly the border of the image.
Sort by magnitude. The gradient isn't just a direction; the relative sizes tell you which knobs give the most bang for the buck.
117 of the top 130 are last-layer weights: at the start, the knobs closest to the output matter most.
Intuition: the gradient ranks the knobs by how much the cost cares about each.
If asked: "Why so tiny in layer 1?" Each first-layer weight's effect is diluted through two squishing layers. That's the vanishing-gradient effect from the ReLU slide, in miniature.
Gradient descent · our training run
Repeat: compute the gradient, step, repeat
With all 60,000 training images, 3Blue1Brown's version reaches ~96% (~98% with tweaks); the best MNIST models, over 99.7%.
Our run: 8,500 MNIST digits, batches of 32, SGD + momentum, 60 epochs, 1,500 held-out test digits. Reference figures: 3Blue1Brown ch. 2 (96%, 98%); MNIST records ≈99.8%.
~1.5 min
Left: all 13,002 parameters of our network as colours. The 16 tiles are the first-layer weight images (one 28×28 tile per hidden neuron); small panels are layers 2 and 3 and the biases. Right: cost and test accuracy.
After 1 epoch (one pass over the data). Cost already fell a lot, but accuracy is still poor.
Epoch 3: it "clicks"; accuracy jumps.
Epoch 10: about 91%.
Epoch 60: 94.1% on digits it never saw. The tiles drift from random static to structured blobs.
Context: more data and better networks go much further. Each tile colour scale is normalised per frame.
Intuition: training is thousands of tiny downhill steps; the weights slowly organise themselves.
If asked: "Why did cost drop before accuracy?" Early on the network learns to output ~0 everywhere (nine of ten targets are 0), which lowers squared error without getting answers right.
Inside the trained network
What did the hidden neurons actually learn?
What we hoped: tidy edge detectors. (Our drawing.)
What it learned: our network's real first-layer weights, one tile per neuron.
Loose blobs of + and −, mostly noise-like. Still 94% accurate.
It found weights that work , not weights that match our story.
~1.5 min
Back to the hope from Part 1. If the first layer were edge detectors, its weight images would look like these little oriented strokes. These are drawn by us, not learned.
Now the real thing: each tile is one first-layer neuron's 784 weights reshaped to 28×28. Blue positive, rose negative.
There's some structure (blobs, a few stroke-ish shapes) but nothing like clean edges. Yet the network gets 94% right.
Gradient descent found a local minimum that solves the training task. Nothing forced it to be human-interpretable.
Intuition: the optimiser only cares about the cost, not about our explanations.
If asked: "So is the layered story wrong?" For tiny MLPs, mostly. For big convolutional networks trained on natural images, early layers really do learn edge and colour detectors. Interpretability is its own research field (think mechanistic interpretability).
Inside the trained network
Feed it pure noise and it confidently says "0"
Of 500 random-noise images, 83% got a top score above 0.5. Every one was called a 0 or a 3.
It has only ever seen digits, so it has no way to say "none of these".
~1 min
Random pixels: uniformly random brightness everywhere. Obviously not a digit.
Run it through the real network.
Output "0" at about 0.9. It's sure.
Not a fluke: we tried 500 noise images. Most got a confident answer, always 0 or 3.
Its whole universe is centred 28×28 digits. Nothing in training ever rewarded "I don't know", so the output can't express it.
Intuition: a network only knows the world its training data showed it.
Analogy: a coin sorter will happily classify a washer as a quarter.
If asked: "How do real systems handle this?" Out-of-distribution detection, calibration, and adding "none of the above" examples to training. It's still an open problem, and it's exactly what our project measures: how models behave on combinations they never saw.
Inside the trained network
What this tells us
Local minimum
It solved the task its own way. Good accuracy ≠ human-like features.
Narrow world
It's only been graded on centred digits; anything else gets a confident guess.
Bigger nets do better
Convolutional nets learn edge-like filters early and parts later.
Memorisation
Big nets can even fit random labels. Structure in the data is what makes learning generalise.
Zeiler & Fergus, "Visualizing and Understanding Convolutional Networks" (ECCV 2014); Olah et al., "Feature Visualization" (Distill, 2017); Zhang et al., "Understanding deep learning requires rethinking generalization" (ICLR 2017).
~1.5 min
Gradient descent finds a good-enough setting, not the one we imagined.
Its training data is its whole world. This is why we obsess over held-out tests in our project.
With more structure (convolutions) and more data, learned features do become interpretable: Zeiler & Fergus visualised edge and texture detectors in early layers, parts and objects deeper. Distill's Feature Visualization shows the same across a big image model.
Zhang et al. trained standard image networks on images with shuffled, random labels. They still hit ~100% training accuracy, by memorising. With real labels they generalise. So "it can fit the data" says little; what matters is what it learns from structure.
Intuition: networks are lazy optimisers: they learn whatever lowers the cost, which may or may not be the "real" concept.
If asked: "Does this mean MLPs are useless?" No: the MLP blocks inside transformers (deck 3) are exactly these layers, just huge. The lesson is about what the objective does and doesn't force.
PART 3
Backpropagation: who should change, and by how much
The algorithm that computes the gradient, one example at a time.
~20 s
We know we want . Backprop is how we compute it efficiently. First intuition, then (Part 4) the calculus.
Backpropagation · one example
One example: which outputs should move, and how far?
Target: 1 for "2" , 0 for the rest.
Nudges proportional to how far off each output is.
Focus on one: how do we raise the "2" neuron?
~1 min
Back to the untrained network (our real starting weights) and one training example: a 2.
We want the "2" output to go to 1 and the others to 0.
The amber arrows are the nudges we'd like: up for "2", down for the rest, longer when further from target: the arrow runs from to . Horizontal here, but read them as "push up" (toward 1) or "push down" (toward 0).
We can't set outputs directly; we can only change weights and biases. So: what can make the "2" neuron brighter?
Intuition: backprop starts with "what do I want the outputs to do?" and works backwards.
If asked: "Why proportional to the error?" Because that's what the derivative of squared error says: . Bigger miss, bigger push.
Backpropagation · three levers
Three levers raise one neuron
1 · Raise the bias .
2 · Change in proportion to : bright neurons' weights matter most. Fire together, wire together.
3 · Change in proportion to . We can't set it directly, so it becomes a wish for the layer before.
~2 min
The "2" neuron's activation is sigmoid of a weighted sum of the 16 neurons before it, plus a bias. Real activations and weights shown (untrained net, this 2).
Lever 1, the bias: raising it raises the output, full stop.
Lever 2, the weights. A weight's effect is multiplied by the activation it carries: a dark neuron (0) contributes nothing however you set its weight. So change weights in proportion to the activations. Hebb's rule from neuroscience: neurons that fire together wire together, which is what this looks like.
Lever 3, the previous activations: increase neurons with positive weights, decrease ones with negative weights, in proportion to the weight. We can't touch activations directly, but we can record this as a wish.
Intuition: each lever's sensitivity is "the thing it multiplies".
If asked: "Is this how brains learn?" Only loosely. The Hebbian echo is real, but brains don't obviously run exact backprop; that's an active research question.
Backpropagation · going backwards
Add up every output's wishes and pass them back
Each output wants to move toward its target.
Each layer-2 neuron sums the wishes sent to it, weighted by its connections.
…and passes its own wishes one layer further back.
Along the way, every weight and bias gets its nudge for this example.
Repeat for every example and average: that average is .
~1.5 min
Whole network, untrained, looking at our 2. All the values are computed live from real weights.
Ten outputs, ten wishes (amber arrows: up or down, length = size).
Lever 3 from every output at once: each layer-2 neuron gets ten requests and adds them up. Amber lines = how much each connection carries.
Now layer 2 has wishes of its own, so repeat: pass them to layer 1. Same rule, one layer back. This recursion is the "back-propagation".
At every layer, levers 1 and 2 turn the wishes into nudges for that layer's biases and weights, all the way to the 12,544 input weights.
That's the nudge from one example. A 2 alone would teach "everything is a 2", so we average the nudges over all examples. The average is the negative gradient.
Intuition: blame flows backward, split according to each connection's strength.
If asked: "Why backward and not forward?" Going backward reuses each layer's summed wishes for everything earlier, so one backward pass costs about as much as one forward pass, instead of one pass per parameter.
Backpropagation · in practice
In practice: quick noisy steps with mini-batches
Careful walker: each step averages all 8,500 examples. Accurate, slow.
Stumbler: each step uses a mini-batch of 32. Noisy, 266× cheaper.
Our run: 60 passes × 266 batches ≈ 16,000 steps.
Stochastic gradient descent (SGD)
~1.5 min
Back to the 2-knob landscape.
True gradient descent averages backprop over the whole training set before every step. Each step is excellent and expensive.
Instead: shuffle, chop into mini-batches of 32, take a step per batch. Each step is a noisy estimate of the true gradient, but you take hundreds of them for the same compute. A stumbling walker who moves fast beats a careful one who checks the map before every step.
For our network: 8,500 ÷ 32 ≈ 266 steps per epoch, 60 epochs.
This is SGD. Every big model is trained this way, with fancier step rules (momentum, Adam).
Intuition: a rough direction now beats a perfect direction later.
If asked: "How do you pick the batch size?" Mostly by hardware: as big as fits on the GPU efficiently, then tune the learning rate with it. Deck 7 covers DataLoaders and batching.
PART 4
The calculus: chain rule all the way down
Same story as Part 3, now as exact derivatives.
~20 s
If the last section made sense, this is the same thing written precisely. If calculus is rusty: it's just the chain rule.
Backprop calculus · simplest case
The simplest network: one neuron per layer
Nudge a little: how much does change? That ratio is .
~1.5 min
Strip the network to one neuron per layer, so each layer has one weight and one bias. Superscripts are layer indices, not powers.
Focus on the last connection: from to .
The cost for one example: squared distance to the target .
Name the weighted sum before the squish: . This intermediate is the key bookkeeping trick.
Then is just sigmoid of .
Watch: bump from 1.20 to 1.30; , then , then the cost all shift. The derivative is the ratio of the cost's change to the weight's change, for a tiny bump.
Intuition: a nudge to causes a nudge to , which causes a nudge to , which causes a nudge to . A chain of causes.
If asked: "Why the subscript 0 on ?" It's the cost of training example number 0; the full cost averages over all of them.
Backprop calculus · chain rule
Multiply the sensitivities along the path
~2.5 min
The same chain as a computational graph: , and feed ; feeds ; and feed the cost.
The chain rule: the total sensitivity is the product of the sensitivities along the path. A nudge passes through three links.
Link 1: , so . The weight's influence is scaled by the activation it multiplies: lever 2 from Part 3.
Link 2: derivative of sigmoid at . Near 0 in the tails: the vanishing-gradient point.
Link 3: derivative of is : the error itself.
Put together: input activity × sensitivity × error. "Fire together, wire together", made exact.
For the bias, the first link is just 1. Lever 1.
For the previous activation, the first link is the weight. Lever 3. This quantity is what we feed into the same formulas one layer back, and so on to the input: that's backpropagation.
Everything so far is for one example. The full gradient averages over the training examples (or the mini-batch).
Intuition: sensitivity of a chain = product of each link's sensitivity.
If asked: "What's ?" , which you can compute from for free: . That's why we cache activations on the forward pass.
Backprop calculus · many neurons
More neurons: same rule, plus a sum over paths
That sum is Part 3's "add up every output's wishes".
~2 min
Now several neurons per layer. Only the bookkeeping changes.
Index neurons in layer by and in layer by . The weight from to is (destination first, matching the row/column of ).
Weighted sum for neuron : the row- dot product, plus its bias.
Cost: sum of squared errors over the outputs.
The new thing: neuron feeds every neuron in layer , so it affects the cost through several paths.
So its derivative is a sum over those paths, each path being the same three-link chain from before. The derivatives for and look exactly like the one-neuron case.
That sum is the "each neuron adds up the wishes sent to it" step from Part 3. Intuition and calculus agree.
Intuition: if you influence the outcome through several channels, your total influence is the sum of the channels.
If asked: "In matrix form?" The vector of these derivatives is times the vector of : backprop is a transposed matrix–vector product per layer.
Bridge forward
Every model today trains with this same loop
model = nn.Sequential( nn.Linear(784 , 16 ), nn.Sigmoid(), nn.Linear(16 , 16 ), nn.Sigmoid(), nn.Linear(16 , 10 ), nn.Sigmoid()) opt = torch.optim.SGD(model.parameters(), lr=0.1 , momentum=0.9 ) for x, y in loader: # mini-batch of 32 a = model(x) # forward pass loss = ((a - y)**2 ).sum(1 ).mean() # cost opt.zero_grad() loss.backward() # backprop opt.step() # gradient step
Transformers (deck 3) and our diffusion U-Net (deck 5) train with this loop. Deck 7 opens up .backward() .
~1.5 min
Our whole network in PyTorch: three linear layers with sigmoids. nn.Linear holds a and a . Six lines.
Forward pass: three times.
The cost: squared error, averaged over the mini-batch.
loss.backward() is backpropagation. PyTorch recorded every operation in the forward pass and walks that graph in reverse, applying the chain rule at each node. That's autograd. You never write the derivatives yourself.
opt.step(): one gradient descent step (with momentum) on all 13,002 parameters.
Round and round: our run did about 16,000 trips. GPT-scale models and our U-Net do the same loop, just with different models, costs and far more steps.
Intuition: architecture changes; the loop forward → cost → backward → step never does.
If asked: "Why zero_grad?" PyTorch accumulates gradients into .grad by default (handy for big effective batches), so you clear them before each backward pass.
Go deeper
Watch these: 3Blue1Brown, Neural Networks ch. 1–4
Chapter 1
Neurons, layers, weights, sigmoid, the matrix form.
Chapter 2
Cost, gradients, and what the network really learned.
Chapter 3
The three levers, nudges flowing backward, SGD.
Chapter 4
The chain rule, term by term.
By Grant Sanderson. Chapters 5–7 (transformers, attention) are deck 3's storyline.
~30 s
This deck followed Grant Sanderson's arc; his videos are the best visual treatment of this anywhere. Each has a written lesson too.
Ch. 1: structure. About 20 minutes.
Ch. 2: gradient descent, plus the "what did it learn" analysis.
Ch. 3: backprop intuition.
Ch. 4: the calculus. Watch it twice.
Next deck continues with his transformer chapters.
If asked: "Where's the code?" Michael Nielsen's free online book Neural Networks and Deep Learning builds a closely related MNIST network (784-30-10) in Python, and 3Blue1Brown's series draws on it.
Recap
Neural networks in six lines
A neuron holds a number; a layer is .
The network is one function with 13,002 knobs.
The cost turns "how wrong" into one number.
Gradient descent walks the knobs downhill, in noisy mini-batch steps.
Backprop computes the gradient: the chain rule, layer by layer, in reverse.
What it learns is whatever lowers the cost, not necessarily our story.
~1 min
Structure: neurons, weighted sums, squish.
Parameters: everything the network "knows" lives in these numbers.
Objective: the cost.
Optimiser: SGD.
Gradient computation: backprop = chain rule, which autograd automates.
The humbling lesson: good accuracy, alien internals, overconfident on garbage. Carry this into the project.
Intuition: deck 3 swaps the structure (attention instead of a plain MLP) and keeps everything else.
If asked: mini-presentation idea: "the gradient of one neuron, derived live on the whiteboard".