QMIND Research · Deck 3 of 7

Transformers, from the residual stream up

What a GPT actually computes, one vector at a time.

Follows the arc of 3Blue1Brown's Deep Learning chapters 5–7 (Grant Sanderson). The words and visuals here are our own.

Context · What an LLM does

A language model just predicts the next token

Predict a distribution, sample one token, append it, repeat. That loop is all “generation” is.

Context · Why sequences were hard

Before 2017, models read one word at a time

Sequential. One word per step, left to right.

Bottleneck. Everything so far squeezed into one fixed-size vector.

Slow. Step waits for step , so the GPU idles.

Context · The 2017 idea

2017: every word looks at every other, at once

Parallel

All positions update in the same step: a few big matrix multiplies. Exactly what GPUs are built for.

Attention Is All You Need

Vaswani et al., Google, NIPS 2017. Built for translation; trained in 3.5 days on 8 GPUs.

Source: Vaswani et al., “Attention Is All You Need”, NIPS 2017 (research.google/pubs/attention-is-all-you-need)

Context · The name

GPT: Generative Pre-trained Transformer

G

Generative

It writes: predicts a distribution, samples, repeats.

P

Pre-trained

First learns from a huge pile of text, then gets adapted.

T

Transformer

The 2017 architecture. The rest of this deck.

GPT-1Jun 2018 GPT-2Feb 2019 · 1.5B params GPT-3May 2020 · 175B paramsour reference model today ChatGPTNov 2022

Sources: Brown et al., “Language Models are Few-Shot Learners” (2020); Wikipedia, “Generative pre-trained transformer” (release dates)

Context · The map

The whole architecture fits on one slide

PART B

Into the residual stream

Text becomes tokens, tokens become vectors, and the vectors flow down a stream that every layer edits.

Residual stream · Tokens

The model reads integers, not letters

GPT-3’s vocabulary: 50,257 tokens. Whole words, word pieces, punctuation.

One ID. The three r’s are not in the input at all.

Token IDs: GPT-2/GPT-3 BPE tokenizer (r50k_base), computed with the gpt-tokenizer library

Residual stream · Embedding

The embedding matrix turns each ID into a vector

Dimensions: Brown et al. 2020, Table 2.1 (GPT-3 175B: d_model = 12,288); vocabulary from the GPT-2 tokenizer

Residual stream · Meaning as direction

Directions in embedding space carry meaning

Each token is a point in a 12,288-dimensional space. We can only draw three.

Same offset, same meaning: a “gender” direction.

Another direction: plural.

Nobody designed these directions. They fall out of training.

Residual stream · Measuring alignment

Dot products measure how much two vectors agree

+0.98

In 12,288-D: “how far does this vector point along the plural direction?”

Residual stream · The backbone

Every layer reads the stream and adds to it

Add, never overwrite.

Residual stream · Width and order

A fixed context window, with positions added in

GPT-3 context: Brown et al. 2020 (n_ctx = 2,048). Larger windows: Meta Llama 3.1 (Jul 2024, 128K); Google Gemini 1.5 Pro (Feb 2024, up to 1M)

PART C

Attention: vectors talking to each other

How information moves between positions, so that each word’s vector can soak up its context.

Attention · Why we need it

Context decides what a word means

After the lookup, every “mole” gets the same vector.

The mole dug a tunnel under the lawn.

One mole of gas is molecules.

The doctor checked the mole on her arm.

The agency suspected a mole in the embassy.

Attention · Queries and keys

Queries ask questions, keys advertise answers

dimensions

Attention · The pattern

Score every key against every query

Each cell: . Blue positive, rose negative.

Big scores = “this word is relevant to that one”.

softmax down each column

Scale the scores up: sharper, one winner.

The model sets its own sharpness through and .

Attention · Masking

Causal masking: no peeking at the future

A key that comes after the query: score set to .

After softmax those weights are exactly 0.

So every column is a training example:

a → fluffy
a fluffy → blue
a fluffy blue → creature
… 7 examples, one pass

Attention · The formula

The whole formula, one term at a time

Every query dotted with every key: the whole score grid in one matrix multiply.

Dot products of random 128-D vectors are large: spread about .

Divide by : spread about 1, so softmax stays soft and gradients survive.

Normalize into attention weights.

Use the weights to mix the value vectors. Next slide.

Attention · Values

Values carry the message that gets added

Attention · Value map

The value map squeezes through 128 dimensions

full map: M weights
M
value-down
value-up
128-D bottleneck

In code, the value-up maps of all heads are packed into one output matrix .

Attention · Multi-head

Many heads, many jobs, all at once

Head types: Olsson et al., “In-context Learning and Induction Heads” (Anthropic, 2022); Clark et al., “What Does BERT Look At?” (2019). Patterns drawn here are illustrative.

Attention · By the numbers

Counting GPT-3’s attention parameters

PieceShapeWeights
Query , key , value-down, value-up6,291,456
One layer: 96 heads603,979,776
All 96 layers57,982,058,496

About 58 billion of GPT-3’s 175 billion: about one third.

Sources: Brown et al. 2020, Table 2.1 (96 layers, d_model 12,288, 96 heads, d_head 128); 3Blue1Brown, “Attention in transformers” (lesson text)

Attention · The catch

Attention’s cost grows with the square of context

Double the context, four times the scores.

GPT-3 at 2,048 tokens: M scores per head, per layer.

Deck 4: why this dominates long-context serving, and how the KV cache helps.

PART D

MLPs, facts, and the geometry of high dimensions

Two thirds of the weights, and a surprising amount of room.

MLP & facts · The block

The MLP processes each vector on its own

: ReLU or GELU

Same MLP at every position, separately. No mixing between tokens.

MLP & facts · A toy lookup

A toy fact lookup: Gretzky plays hockey

Toy picture. Evidence that MLPs store associations: Geva et al., “Transformer Feed-Forward Layers Are Key-Value Memories” (2021); Meng et al., “Locating and Editing Factual Associations in GPT” (2022)

MLP & facts · By the numbers

Two thirds of the parameters live in the MLPs

Sources: Brown et al. 2020, Table 2.1; 3Blue1Brown, “How might LLMs store facts” (lesson text). Biases and layer norms omitted.

MLP & facts · High dimensions

High-dimensional space is roomier than it looks

Random vectors: has spread .

Johnson–Lindenstrauss: the number of nearly perpendicular directions grows exponentially with .

MLP & facts · Superposition

Superposition: more features than neurons

Elhage et al., “Toy Models of Superposition” (2022); Bricken et al., “Towards Monosemanticity” (2023); Templeton et al., “Scaling Monosemanticity” (2024), all Anthropic

MLP & facts · How the pieces divide the work

Attention moves information, MLPs transform it

PART E

From the last vector to the next word

Unembedding, softmax, temperature, training, and the whole stack in one pass.

Putting it together · Output

Unembedding: the last vector scores every token

: one row per token
Putting it together · Sampling

Temperature controls how adventurous sampling is

: the model’s own distribution.

: sharper. Safer, more repetitive.

: always the top token (greedy).

: flatter. More surprising, more mistakes.

Putting it together · Training

Training is deck 2’s loop on next-token guesses

loss

GPT-3: 300 billion training tokens. Llama 3: over 15 trillion.

Sources: Brown et al. 2020 (300B tokens); Meta Llama 3 announcement via The Register, Apr 2024 (15T+ tokens)

Putting it together · End to end

The whole stack, end to end

Putting it together · Recap

Six ideas to carry into deck 4

1 · Tokens

Text → integer IDs

The model never sees letters.

2 · Embeddings

ID → vector

Directions in 12,288-D carry meaning.

3 · Residual stream

One vector per position

Every block reads it and adds an edit.

4 · Attention

Moves information

Queries match keys; values get added. Many heads in parallel.

5 · MLPs

Process in place

Two thirds of the weights; superposition packs in features.

6 · Output

Logits → softmax → sample

Trained by gradient descent on next-token prediction.

Putting it together · Next

Watch these, then let’s scale it up

3Blue1Brown, Deep Learning (Grant Sanderson)

Deck 4: serving, scale, RL and test-time compute.

Elapsed
0:00:00
Steps on this slide
Next slide