QMIND Research · Deck 2 of 7

Neural networks from scratch

A handwritten digit, 13,002 numbers, and the loop that tunes them.

The storyline follows 3Blue1Brown's Neural Networks series, chapters 1–4, by Grant Sanderson. Words and visuals here are ours; the videos are linked at the end.

The puzzle

You read a sloppy 3 instantly. Could you code it?

To a computer: three grids of 784 numbers that barely agree.

So we won't write rules. We'll tune a function on examples.

Roadmap

Four ideas, one tiny network

Part 1

Structure

Neurons, layers, weights, biases. The network is a function.

Part 2

Gradient descent

Score it with a cost; walk downhill. What did it learn?

Part 3

Backpropagation

Which knobs to turn, and how much, for one example.

Part 4

The calculus

The chain rule that makes Part 3 exact.

Same order as 3Blue1Brown chapters 1–4. Links on the last slides.

PART 1

What is a neural network?

Neurons that hold numbers, layers that feed each other, and 13,002 knobs.

Structure

A neuron is just a thing that holds a number

Its activation: how lit up it is.

0 = dark, 1 = fully on.

Everything else is rules for computing these numbers.

Structure · input layer

A 28×28 image becomes 784 input neurons

28 × 28 = 784 pixels, each a brightness from 0 to 1

Unroll them into one column

One neuron per pixel

Structure · layers

Two hidden layers of 16, then ten outputs

10 outputs, one per digit: how strongly it thinks "this is a 7".

2 hidden layers of 16. An arbitrary choice: small enough to watch.

Every neuron feeds every neuron in the next layer.

The answer = the brightest output.

Structure · forward pass

Each layer's activations determine the next

Real network, real weights. Brighter line = bigger weight times activation. Blue pushes the next neuron up, rose pushes it down.

Structure · the hope

The hope: layers build edges into parts into digits

Reusable parts: an 8 is two loops, a 6 is a loop and a line.

Spoiler: our trained net does not do this cleanly
Structure · one neuron

One neuron: weigh every pixel, then add

Positive weights where you want ink, negative around it: a detector for a horizontal stroke.

Structure · squish

Squish the sum into (0, 1); the bias sets the bar

Very negative → near 0. Very positive → near 1.

Bias : how big the sum must be before the neuron lights up.

Structure · modern default

Modern networks mostly use ReLU instead

Sigmoid's slope is at most and nearly 0 in the tails, so learning signals fade layer after layer.

ReLU's slope is 1 wherever it's on. Today's nets use it or smooth cousins (GELU, SiLU). We keep sigmoid to match the classic.

Structure · in matrix form

A whole layer is one matrix–vector product

Toy numbers. The real first layer: is 16 × 784. GPUs are built to do exactly this.

Structure · count the knobs

Count the knobs: 13,002 parameters

weights
0
weights
0
weights
0
biases
0
total
0

The whole network is one function with 13,002 knobs .

PART 2

How a network learns: gradient descent

Score how wrong it is. Then walk downhill.

Gradient descent · cost

Score one guess: the cost of a single example

Trained: small cost means confident and correct.

Gradient descent · cost

Average the cost over every training example

Network

784 pixels in → 10 numbers out

Cost

13,002 knobs in → 1 number out

Learning = find knobs that make it small.

Gradient descent · one knob

Minimise by walking downhill

Slope says which way is down.

Steps shrink as it flattens out.

Different start, different valley. No guarantee of the best one.

Gradient descent · two knobs

With two knobs: step against the gradient

points uphill; is the steepest way down.

Small steps: steady progress.

Too big: overshoot and zigzag.

Same idea with 13,002 knobs. We just can't draw it.

Gradient descent · the gradient

: 13,002 nudges, and some matter far more

Our real gradient at the start of training. Blue: turn this knob up. Rose: turn it down.

32% are exactly zero: weights on pixels that were blank in every image of this batch.

Sorted by size: a few nudges dwarf the rest.

Top 1% (130 entries) carry 40% of the total size. The biggest is ~320× the median nonzero nudge.

Gradient descent · our training run

Repeat: compute the gradient, step, repeat

With all 60,000 training images, 3Blue1Brown's version reaches ~96% (~98% with tweaks); the best MNIST models, over 99.7%.

Our run: 8,500 MNIST digits, batches of 32, SGD + momentum, 60 epochs, 1,500 held-out test digits. Reference figures: 3Blue1Brown ch. 2 (96%, 98%); MNIST records ≈99.8%.

Inside the trained network

What did the hidden neurons actually learn?

What we hoped: tidy edge detectors. (Our drawing.)

What it learned: our network's real first-layer weights, one tile per neuron.

Loose blobs of + and −, mostly noise-like. Still 94% accurate.

It found weights that work, not weights that match our story.

Inside the trained network

Feed it pure noise and it confidently says "0"

Of 500 random-noise images, 83% got a top score above 0.5. Every one was called a 0 or a 3.

It has only ever seen digits, so it has no way to say "none of these".

Inside the trained network

What this tells us

Local minimum

It solved the task its own way. Good accuracy ≠ human-like features.

Narrow world

It's only been graded on centred digits; anything else gets a confident guess.

Bigger nets do better

Convolutional nets learn edge-like filters early and parts later.

Memorisation

Big nets can even fit random labels. Structure in the data is what makes learning generalise.

Zeiler & Fergus, "Visualizing and Understanding Convolutional Networks" (ECCV 2014); Olah et al., "Feature Visualization" (Distill, 2017); Zhang et al., "Understanding deep learning requires rethinking generalization" (ICLR 2017).

PART 3

Backpropagation: who should change, and by how much

The algorithm that computes the gradient, one example at a time.

Backpropagation · one example

One example: which outputs should move, and how far?

Target: 1 for "2", 0 for the rest.

Nudges proportional to how far off each output is.

Focus on one: how do we raise the "2" neuron?

Backpropagation · three levers

Three levers raise one neuron

1 · Raise the bias .

2 · Change in proportion to : bright neurons' weights matter most. Fire together, wire together.

3 · Change in proportion to . We can't set it directly, so it becomes a wish for the layer before.

Backpropagation · going backwards

Add up every output's wishes and pass them back

Each output wants to move toward its target.

Each layer-2 neuron sums the wishes sent to it, weighted by its connections.

…and passes its own wishes one layer further back.

Along the way, every weight and bias gets its nudge for this example.

Repeat for every example and average: that average is .

Backpropagation · in practice

In practice: quick noisy steps with mini-batches

Careful walker: each step averages all 8,500 examples. Accurate, slow.

Stumbler: each step uses a mini-batch of 32. Noisy, 266× cheaper.

Our run: 60 passes × 266 batches ≈ 16,000 steps.

Stochastic gradient descent (SGD)

PART 4

The calculus: chain rule all the way down

Same story as Part 3, now as exact derivatives.

Backprop calculus · simplest case

The simplest network: one neuron per layer

Cost
Weighted sum
Squish

Nudge a little: how much does change? That ratio is .

Backprop calculus · chain rule

Multiply the sensitivities along the path

Backprop calculus · many neurons

More neurons: same rule, plus a sum over paths

That sum is Part 3's "add up every output's wishes".

Bridge forward

Every model today trains with this same loop

model = nn.Sequential(    nn.Linear(784, 16), nn.Sigmoid(),    nn.Linear(16, 16),  nn.Sigmoid(),    nn.Linear(16, 10),  nn.Sigmoid())opt = torch.optim.SGD(model.parameters(), lr=0.1, momentum=0.9) for x, y in loader:              # mini-batch of 32    a = model(x)                   # forward pass    loss = ((a - y)**2).sum(1).mean() # cost    opt.zero_grad()    loss.backward()                # backprop    opt.step()                     # gradient step

Transformers (deck 3) and our diffusion U-Net (deck 5) train with this loop. Deck 7 opens up .backward().

Go deeper

Watch these: 3Blue1Brown, Neural Networks ch. 1–4

Chapter 1

But what is a neural network?

Neurons, layers, weights, sigmoid, the matrix form.

Chapter 2

Gradient descent, how neural networks learn

Cost, gradients, and what the network really learned.

Chapter 3

What is backpropagation really doing?

The three levers, nudges flowing backward, SGD.

Chapter 4

Backpropagation calculus

The chain rule, term by term.

By Grant Sanderson. Chapters 5–7 (transformers, attention) are deck 3's storyline.

Recap

Neural networks in six lines

  1. A neuron holds a number; a layer is .
  2. The network is one function with 13,002 knobs.
  3. The cost turns "how wrong" into one number.
  4. Gradient descent walks the knobs downhill, in noisy mini-batch steps.
  5. Backprop computes the gradient: the chain rule, layer by layer, in reverse.
  6. What it learns is whatever lowers the cost, not necessarily our story.
Elapsed
0:00:00
Steps on this slide
Next slide