QMIND Research · Deck 3 of 7
Transformers, from the residual stream up
What a GPT actually computes, one vector at a time.
Follows the arc of 3Blue1Brown's Deep Learning chapters 5–7 (Grant Sanderson). The words and visuals here are our own.
~0.5 min
Deck 2 ended with a 784–16–16–10 network reading digits. Same ingredients today: matrix multiplies, a nonlinearity, gradient descent. Different data: text.
Credit: the arc follows 3Blue1Brown's chapters 5–7. If you liked deck 2, those three videos are the best 75 minutes you can spend after today.
Intuition: a transformer is a long assembly line of vectors; each station reads them and adds a small correction.
If asked: "Is this the same as ChatGPT?" The core network, yes. ChatGPT adds post-training (RLHF and friends), which is deck 4.
Context · What an LLM does
A language model just predicts the next token
Predict a distribution, sample one token, append it, repeat. That loop is all “generation” is.
~1.5 min
The whole job of the network: given text so far, output a probability for every possible next token.
One forward pass gives this distribution. Numbers here are illustrative, not from a real model.
We sample : roll a weighted die. “roamed” wins this time even though “was” was likelier.
Append it and run the whole network again on the longer text. New context, new distribution.
Sample again: “the”.
Sometimes the die picks a long shot (“verdant”, 2%). That randomness is why the same prompt gives different answers.
That is all chat is: this loop, one token at a time, until a stop token.
Intuition: a fantastically good autocomplete, run in a loop.
Analogy: improv storytelling where each person adds one word, except every “person” is the same network rereading the whole story.
If asked: “Does it plan ahead?” Each step only outputs one token, but the internal vectors can encode where the text is going; interpretability work has found evidence of some lookahead. The output interface is still one token at a time.
Context · Why sequences were hard
Before 2017, models read one word at a time
Sequential. One word per step, left to right.
Bottleneck. Everything so far squeezed into one fixed-size vector.
Slow. Step waits for step , so the GPU idles.
~1.5 min
The previous generation: recurrent networks (RNNs, LSTMs). One cell, applied over and over.
Each step takes one word plus the previous hidden state and produces a new hidden state. Watch it crawl.
By “forest”, everything about “fluffy” has to have survived seven squeezes through the same small vector. Old information fades.
Worse for engineering: step 8 cannot start until step 7 is done. You cannot spread one sentence across thousands of GPU cores.
Intuition: an RNN is a game of telephone with itself.
Analogy: reading a book through a straw while only allowed to keep one index card of notes.
If asked: “Didn’t attention exist before 2017?” Yes: Bahdanau et al. (2014) bolted attention onto RNN translators. The 2017 move was deleting the recurrence entirely and keeping only attention plus MLPs.
Context · The 2017 idea
2017: every word looks at every other, at once
Parallel
All positions update in the same step: a few big matrix multiplies. Exactly what GPUs are built for.
Attention Is All You Need
Vaswani et al., Google, NIPS 2017. Built for translation; trained in 3.5 days on 8 GPUs.
Source: Vaswani et al., “Attention Is All You Need”, NIPS 2017 (research.google/pubs/attention-is-all-you-need)
~1.5 min
Same sentence. New rule: no hidden state passed along a chain.
Instead every position gets a direct line to every other position. Twenty-eight lines for eight words.
The model learns how much each line matters. For “creature”, the lines to “fluffy” and “blue” light up. That is attention, which is Part C.
Nothing waits on anything else, so a whole sentence (or 2,048 tokens) is processed in one shot. Training became massively parallel.
The paper: eight Google authors, 2017. Original target was English–German translation, reached state of the art in 3.5 days on 8 GPUs.
Intuition: replace a chain with a fully connected meeting where everyone can hear everyone.
If asked: “Original transformer = GPT?” Not quite. The original had an encoder and a decoder for translation. GPT keeps only the decoder half: a stack that predicts the next token.
Context · The name
GPT: Generative Pre-trained Transformer
G
Generative It writes: predicts a distribution, samples, repeats.
P
Pre-trained First learns from a huge pile of text, then gets adapted.
T
Transformer The 2017 architecture. The rest of this deck.
GPT-1 Jun 2018
GPT-2 Feb 2019 · 1.5B params
GPT-3 May 2020 · 175B params our reference model today
ChatGPT Nov 2022
Sources: Brown et al., “Language Models are Few-Shot Learners” (2020); Wikipedia, “Generative pre-trained transformer” (release dates)
~1 min
Generative: it produces text, using the loop from two slides ago.
Pre-trained: step one is next-token prediction on a giant text corpus. Fine-tuning and RLHF come after (deck 4).
Transformer: the architecture. Everything from here on.
Lineage: GPT-1 (2018), GPT-2 (2019, 1.5B parameters), GPT-3 (2020, 175B), ChatGPT (Nov 2022, built on GPT-3.5).
GPT-3 is our reference all deck because its paper publishes every dimension: 96 layers, width 12,288, 96 heads. Newer frontier models don’t publish these.
Intuition: the architecture barely changed from GPT-2 to GPT-3; the scale changed by ~100×.
If asked: “What about BERT?” Also a transformer (2018, Google), but encoder-only and trained to fill in blanks rather than continue text, so it is not generative in this sense.
Context · The map
The whole architecture fits on one slide
~1.5 min
Here is the full map. Everything after this slide zooms into one box.
Text is chopped into tokens, each an integer ID.
Embedding: each ID looks up a vector, a long list of numbers.
Attention block: vectors exchange information with each other.
MLP block: each vector is processed on its own. Attention + MLP is one layer; GPT-3 stacks 96 of them.
At the end, the last vector is turned into a score for every token in the vocabulary, then probabilities.
Underneath it all runs the residual stream : the vectors themselves, flowing left to right, with every block adding to them. That is Part B.
Intuition: embed, then (talk, think) × 96, then vote on the next token.
If asked: “Where are the 175B parameters?” About one third in attention, two thirds in the MLPs, under 1% in the embeddings. We count them exactly later.
PART B
Into the residual stream
Text becomes tokens, tokens become vectors, and the vectors flow down a stream that every layer edits.
Residual stream · Tokens
The model reads integers, not letters
GPT-3’s vocabulary: 50,257 tokens. Whole words, word pieces, punctuation.
One ID. The three r’s are not in the input at all.
Token IDs: GPT-2/GPT-3 BPE tokenizer (r50k_base), computed with the gpt-tokenizer library
~2 min
The network can only do arithmetic, so text has to become numbers first.
A tokenizer chops text into chunks. Common words are one token; the leading space is part of the token (shown as a dot).
Each chunk is an integer ID. These are the real GPT-3 IDs. Rare words split: “roamed” is two tokens, “verdant” is three.
GPT-3’s vocabulary has 50,257 entries, learned by byte-pair encoding: start from bytes, repeatedly merge the most frequent pair.
Classic puzzle: “how many r’s in strawberry?” Models used to get this wrong.
In GPT-3’s tokenizer, “ strawberry” is a single ID, 41236. The letters never reach the network. It can only know the spelling if it memorized facts about that ID from training text.
Intuition: the model sees a sequence of opaque symbols, like reading Chinese characters, not an alphabet.
If asked: “Why not just use characters?” Sequences would be ~4× longer and attention cost grows with length squared. Subword tokens are the compromise. Modern tokenizers are larger (Llama 3 has a 128K-token vocabulary).
Residual stream · Embedding
The embedding matrix turns each ID into a vector
Dimensions: Brown et al. 2020, Table 2.1 (GPT-3 175B: d_model = 12,288); vocabulary from the GPT-2 tokenizer
~1.5 min
The embedding matrix : one column per vocabulary token, 50,257 columns. Each column has 12,288 numbers in GPT-3.
Token “ creature” has ID 7185. The ID is just an address: go to column 7185.
Pull that column out. That is the token’s starting vector. No computation, a pure lookup.
Do it for every token. The sentence is now a list of vectors, one per position. This list is where the residual stream starts.
617 million weights, and they are learned by gradient descent like every other weight. Nobody hand-designs what a column means.
Intuition: a giant lookup table from symbol to starting meaning.
Analogy: a dictionary where each entry is written in a 12,288-number language the model invented.
If asked: “Is this a matrix multiply?” Mathematically it is times a one-hot vector; in code it is an index lookup (nn.Embedding), which is much cheaper.
Residual stream · Meaning as direction
Directions in embedding space carry meaning
Each token is a point in a 12,288-dimensional space. We can only draw three.
Same offset, same meaning: a “gender” direction.
Another direction: plural.
Nobody designed these directions. They fall out of training.
~2 min
A cartoon: positions are made up, but the geometry is the real idea.
Every token is a vector: an arrow from the origin, or just its tip. Similar words end up near each other.
The arrow from man to woman is nearly the same as king to queen. So “female-ness” is a direction , not a location.
Rotate: a different direction encodes plural. man→men, cat→cats run parallel.
So you can do arithmetic on meaning: start at king, add the gender offset, land near queen. Classic word2vec result (Mikolov et al., 2013).
The whole deck rests on this: meaning lives in directions, and layers move vectors along them.
Intuition: a vector’s position along a direction answers a yes/no-ish question about the word.
If asked: “Does the analogy really work?” Approximately. The nearest neighbour to is queen once you exclude the input words; in a transformer, these clean directions are clearest in the embedding layer and get richer and more tangled deeper in.
Residual stream · Measuring alignment
Dot products measure how much two vectors agree
In 12,288-D: “how far does this vector point along the plural direction?”
~1 min
One tool we need constantly: the dot product.
Multiply matching coordinates and add. Geometrically, lengths times cosine of the angle. Pointing the same way: large and positive.
Perpendicular: zero. The vectors have nothing to say about each other.
Opposite: negative.
Sweep a full turn and the dot product traces a cosine.
In the model, dotting a vector with a meaningful direction (plural, gender, “is an adjective”) asks how much of that property it has. Attention and MLPs are built out of exactly these questions.
Intuition: dot product = “how much do these agree?”
If asked: “Why not cosine similarity?” Cosine similarity is the dot product with lengths divided out. Inside the network, length carries information too, so raw dot products are used.
Residual stream · The backbone
Every layer reads the stream and adds to it
~2 min
One lane per token position, one vector per lane. These are the embeddings we just looked up.
An attention block reads all lanes, lets them exchange information (fluffy and blue talk to creature), and writes a small change, a delta, back into each lane.
An MLP block works on each lane separately and adds its own delta.
Repeat 96 times in GPT-3. By the end, each vector has been edited by every block. The last lane’s vector is what predicts the next token.
The key equation: each block adds its output to the stream. This is the residual (skip) connection from ResNets (He et al., 2015).
Intuition: the stream is a shared whiteboard; blocks read it and add notes, never erase.
Analogy: a document passed through 192 reviewers, each allowed only to add tracked changes.
If asked: “Why add instead of replace?” Two reasons. Gradients flow straight back through the additions (remember backprop from deck 2), so 96-layer stacks train. And blocks only need to learn small edits. Subtle point: in GPT each block reads a normalized copy, .
Residual stream · Width and order
A fixed context window, with positions added in
GPT-3 context: Brown et al. 2020 (n_ctx = 2,048). Larger windows: Meta Llama 3.1 (Jul 2024, 128K); Google Gemini 1.5 Pro (Feb 2024, up to 1M)
~1.5 min
Two practical facts about the stream.
It has a fixed number of lanes: the context window. GPT-3 had 2,048. Text that scrolls out of the window is simply gone; that is why early ChatGPT “forgot” the start of long chats. Modern models have 128K to 1M+.
Second: attention by itself treats its inputs as an unordered set. “dog bites man” and “man bites dog” would give the same set of vectors.
Fix: add a position vector to each embedding. GPT-2 and GPT-3 learn one vector per position. Now the same word at a different position is a different vector.
Intuition: order is not built in; it is stamped onto each vector.
If asked: “What do modern models use?” Mostly RoPE (rotary position embeddings, Su et al. 2021): instead of adding a vector, rotate each query and key by an angle proportional to position, so the dot product depends on relative distance. Also, the causal mask (Part C) leaks some order information on its own.
PART C
Attention: vectors talking to each other
How information moves between positions, so that each word’s vector can soak up its context.
~0.3 min
Part C: the block that made transformers famous.
One job: move information from one position’s vector into another’s.
Intuition: attention is a learned, content-based routing system.
Attention · Why we need it
Context decides what a word means
After the lookup, every “mole” gets the same vector.
The mole dug a tunnel under the lawn.
One mole of gas is molecules.
The doctor checked the mole on her arm.
The agency suspected a mole in the embassy.
~1 min
The embedding lookup has no context. Token “ mole” gets the identical vector in all four sentences.
Animal sense: words like “dug” and “tunnel” should push the vector toward the burrowing-animal region.
Chemistry sense: “gas” and “molecules” push it somewhere else entirely.
Skin sense: “doctor”, “arm”.
Spy sense: “agency”, “embassy”. Attention’s job is to compute those violet arrows: the context-dependent change to add.
Intuition: attention lets each vector pull in information from the other words so it can mean the right thing.
Analogy: you hear “bank” and wait for “river” or “loan” before deciding which one.
If asked: “Does this only matter for ambiguous words?” No. Every word gets refined: “tower” after “Eiffel” should carry Paris, iron, landmark. Ambiguity is just the clearest demo.
Attention · Queries and keys
Queries ask questions, keys advertise answers
~1.5 min
Running example: “a fluffy blue creature roamed the verdant forest”. Imagine one head whose job is: adjectives update the nouns they describe. (A cartoon; real heads are messier.)
“creature” wants to ask: any adjectives in front of me? That question is a query vector: times its embedding, a much smaller 128-D vector.
The same matrix is applied at every position, so every token produces a query.
A second matrix produces a key per token: what it can offer. Adjectives produce keys that say “I’m an adjective”.
If a key matches a query, their dot product is large. The amber lines: creature’s query matches fluffy’s and blue’s keys; forest’s query matches verdant’s.
Intuition: query = what I’m looking for; key = what I contain.
Analogy: a search engine: your query is matched against every page’s keywords.
If asked: “Who decides that a head looks for adjectives?” Nobody. and are learned; whatever matching rule lowers next-token loss is what emerges.
Attention · The pattern
Score every key against every query
Each cell: . Blue positive, rose negative.
Big scores = “this word is relevant to that one”.
Scale the scores up: sharper, one winner.
The model sets its own sharpness through and .
~1.5 min
Columns are queries (the word asking), rows are keys (the word being asked about). Same convention as 3Blue1Brown.
Fill every cell with the dot product of that key and that query. Raw scores, anywhere from very negative to very positive.
The big ones: fluffy→creature, blue→creature, verdant→forest, creature→roamed (a verb asking for its subject).
We want weights, not scores: apply softmax down each column. Every column becomes non-negative and sums to 1. This grid of weights is the attention pattern .
If the scores were twice as big, softmax would get sharper: nearly all weight on the single best match. This is the temperature idea; we come back to it at the output.
The model controls this itself: larger entries mean larger dot products, so sharper attention.
Intuition: each column is one word’s budget of attention, split among the words it cares about.
If asked: “Rows or columns?” The paper writes with tokens as rows and normalizes each row. Same math, transposed. Don’t let the convention trip you up.
Attention · Masking
Causal masking: no peeking at the future
A key that comes after the query: score set to .
After softmax those weights are exactly 0.
So every column is a training example:
a → fluffy a fluffy → blue a fluffy blue → creature … 7 examples, one pass
~1 min
Same scores. One problem: “a” can see “forest”, which comes later.
During training, the model predicts the next token at every position. If position 3 could see position 4, it could just copy the answer. So every score where the key is later than the query is set to minus infinity.
, so softmax gives those cells exactly zero weight, and the rest of each column renormalizes.
Payoff: a single forward pass over one sentence trains on every prefix at once. A 2,048-token sequence gives 2,048 next-token predictions for the price of one pass.
Intuition: each position may only look backwards, so the model can’t cheat.
If asked: “Is the mask needed at inference?” With a KV cache you only compute the newest position, which naturally sees only the past. But the model was trained masked, so it must run the same way. Deck 4 covers the KV cache.
Attention · The formula
The whole formula, one term at a time
Every query dotted with every key: the whole score grid in one matrix multiply.
Dot products of random 128-D vectors are large: spread about .
Divide by : spread about 1, so softmax stays soft and gradients survive.
Normalize into attention weights.
Use the weights to mix the value vectors. Next slide.
~1.5 min
The famous line from the 2017 paper. We already know every piece.
: stack all queries and keys as matrices; one multiply gives every dot product, the grid from two slides ago.
Why ? If query and key entries are roughly unit-variance, their dot product over dimensions has variance 128, standard deviation about 11. Left: the spread. Right: softmax of such scores is essentially one-hot.
One-hot softmax is flat almost everywhere, so its gradients vanish and learning stalls. Dividing by brings the spread back to about 1.
Softmax turns scores into weights (per column in our picture).
: the value vectors, the content that actually gets moved. That is the next slide.
Intuition: is a thermostat that keeps softmax from saturating at initialization.
If asked: “Is it exactly variance ?” Only if entries are independent with unit variance, which is roughly true at initialization. It is a scaling convention that works, not a law.
Attention · Values
Values carry the message that gets added
~1.5 min
Queries and keys only decide who listens to whom . They don’t say what gets said.
A third matrix, , maps each token’s vector to a value : “if you attend to me, here is what to add to yourself”.
Take creature’s column of attention weights (masked, so only earlier tokens): mostly fluffy and blue. Line thickness = weight.
Scale each value by its weight and add them up. That sum is creature’s update, .
Add it into creature’s lane of the residual stream. Now the vector encodes something like “fluffy blue creature”.
Every position does this simultaneously with its own column of weights.
Intuition: attention output = weighted average of messages from the tokens you care about.
Analogy: a group chat: queries/keys decide whose messages you read, values are the messages.
If asked: “Why separate keys and values?” What makes a token findable (it’s an adjective) is different from what it should contribute (blueness). Separate matrices let the model learn both.
Attention · Value map
The value map squeezes through 128 dimensions
In code, the value-up maps of all heads are packed into one output matrix .
~1 min
A value map from 12,288-D to 12,288-D would be a 151-million-weight square, per head. Way too many.
Instead it is factored: value-down to 128-D, then value-up back to 12,288-D. Drawn to scale, the two thin strips are about 2% of the square.
Flow: the full vector is squeezed to 128 numbers, then expanded back. The map has rank at most 128: each head can only move a small slice of information.
Implementation detail: libraries keep one per head (the “down” part) and concatenate all heads’ “up” parts into a single output matrix . Same math, different bookkeeping.
Intuition: each head gets a narrow pipe, so it must be selective about what it copies.
If asked: “Why 128?” In GPT-3, heads . Splitting the width evenly across heads keeps the total cost of multi-head attention equal to one full-width head.
Attention · Multi-head
Many heads, many jobs, all at once
Head types: Olsson et al., “In-context Learning and Induction Heads” (Anthropic, 2022); Clark et al., “What Does BERT Look At?” (2019). Patterns drawn here are illustrative.
~1.5 min
One head = one set of = one attention pattern. GPT-3 runs 96 heads per layer in parallel.
Head A: previous-token head . Each position just looks one step back. Boring, but a crucial building block.
Head B: a syntax head . The verb looks at its subject, the object looks at the verb. Researchers found heads like this in BERT.
Head C: an induction head . The second “Harry” searches for an earlier “Harry” and attends to whatever came right after it: “Potter”.
All heads run at once on the same input; their outputs are added into the stream together.
Result: the model predicts “Potter” by copying a pattern from earlier in the text. This is the basic mechanism of in-context learning.
Intuition: many small specialists reading the same text for different reasons.
Analogy: an editing team: one checks grammar, one tracks names, one looks for repeats.
If asked: “Are heads really this clean?” Some are; many are messy or do several things. Induction heads are one of the best-documented cases: they form abruptly during training, at the same time in-context learning improves.
Attention · By the numbers
Counting GPT-3’s attention parameters
Piece Shape Weights
Query , key , value-down, value-up 6,291,456
One layer: 96 heads 603,979,776
All 96 layers 57,982,058,496
About 58 billion of GPT-3’s 175 billion: about one third.
Sources: Brown et al. 2020, Table 2.1 (96 layers, d_model 12,288, 96 heads, d_head 128); 3Blue1Brown, “Attention in transformers” (lesson text)
~1 min
One head has four maps, each : query, key, value-down, value-up. About 6.3 million weights.
96 heads per layer: about 604 million per attention block.
96 layers: just under 58 billion.
That is a third of GPT-3. Where is the rest? Mostly in the MLPs, Part D.
Intuition: attention is big, but it is not where most of the weights are.
If asked: “Biases, layer norms?” Ignored here; they are a rounding error (millions, not billions). The count also matches per layer, the formula you see in papers.
Attention · The catch
Attention’s cost grows with the square of context
Double the context, four times the scores.
GPT-3 at 2,048 tokens: M scores per head, per layer.
Deck 4: why this dominates long-context serving, and how the KV cache helps.
~0.7 min
Every query meets every key, so an -token context means an grid.
8 tokens: 64 scores. 16: 256. 32: 1,024. Doubling context quadruples the work and the memory for the grid.
At GPT-3’s 2,048 tokens: about 4.2 million scores, for each of 96 heads in each of 96 layers.
This is the main reason long context is expensive. Just flag it; deck 4 goes deep.
Intuition: everyone-talks-to-everyone meetings get expensive fast.
If asked: “Is the MLP also quadratic?” No, MLPs act per position, so their cost is linear in context length. At short contexts the MLP actually dominates compute.
PART D
MLPs, facts, and the geometry of high dimensions
Two thirds of the weights, and a surprising amount of room.
~0.3 min
Part D: the other half of every layer, the part deck 2 already taught you.
Plus the strangest fact about high-dimensional space, which explains how a model can know so much.
Intuition: attention moves information around; the MLP is where much of the looking-up and computing happens.
MLP & facts · The block
The MLP processes each vector on its own
Same MLP at every position, separately. No mixing between tokens.
~1.2 min
This is deck 2’s network, almost exactly: two matrix multiplies with a nonlinearity between.
Up-projection : 12,288 → 49,152 numbers (4× wider). Each output is one dot product, one “neuron”.
Nonlinearity: negatives are zeroed (ReLU) or nearly zeroed (GELU, what GPT uses). Without it, the two matrices would collapse into one.
Down-projection : back to 12,288. Only active neurons contribute.
The result is a delta, added into the residual stream like everything else.
Crucial: the MLP sees one vector at a time. Tokens don’t talk here. Attention moves information; the MLP processes it in place.
Intuition: a wide bank of feature detectors, each of which, when it fires, writes something back.
If asked: “Why 4× wider?” A convention from the original transformer that stuck. Many modern models use gated variants (SwiGLU) with a different ratio, same idea.
MLP & facts · A toy lookup
A toy fact lookup: Gretzky plays hockey
Toy picture. Evidence that MLPs store associations: Geva et al., “Transformer Feed-Forward Layers Are Key-Value Memories” (2021); Meng et al., “Locating and Editing Factual Associations in GPT” (2022)
~1.5 min
How might an MLP store “Wayne Gretzky plays hockey”? A deliberately cartoonish version with named directions.
Incoming vector at the “Gretzky” position. Attention has already moved “first name Wayne” into it, so it points along both name directions.
One row of is a question : “Wayne + Gretzky?” Its dot product with the vector is 2.
Bias −1 brings it to 1; ReLU keeps it. The neuron only fires if both are present: an AND gate.
The matching column of is the answer : hockey + Edmonton Oilers. Scaled by the neuron’s activation, added to the stream.
Swap in Wayne Rooney: dot product 1, bias brings it to 0, neuron silent. Nothing hockey-related gets added.
Intuition: rows of ask questions, columns of hold answers.
If asked: “Is one fact really one neuron?” Almost never. Real facts seem spread across many neurons and layers (superposition, next slides). Mid-layer MLPs are where editing experiments (ROME) found factual associations most localized.
MLP & facts · By the numbers
Two thirds of the parameters live in the MLPs
Sources: Brown et al. 2020, Table 2.1; 3Blue1Brown, “How might LLMs store facts” (lesson text). Biases and layer norms omitted.
~0.8 min
Attention: 58 billion (previous count).
MLPs: each block has two matrices, 1.2 billion weights. Times 96 layers: 116 billion.
Embedding and unembedding: 0.6 billion each, a sliver. Total: about 175 billion. The count closes.
Per layer: the MLP is twice the size of attention. That ratio holds in most dense transformers.
Intuition: attention routes; the MLPs are the bulk storage and compute.
If asked: “Do embedding and unembedding share weights?” In GPT-2 they are tied (the same matrix, transposed). Counting them separately, as here, is how the total lands on 175B; either way it is under 1%.
MLP & facts · High dimensions
High-dimensional space is roomier than it looks
Random vectors: has spread .
Johnson–Lindenstrauss: the number of nearly perpendicular directions grows exponentially with .
~1.2 min
Experiment, computed live in your browser: 70 random vectors, histogram of all pairwise angles.
In 2-D, angles are all over the place.
10-D: already bunching around 90°.
100-D: almost all within about ±10° of perpendicular. The spread of shrinks like .
1,000-D: within a few degrees.
GPT-3’s 12,288-D: random directions are within about a degree of perpendicular. Exactly perpendicular, you only get directions. Nearly perpendicular, you get exponentially many.
Intuition: a 12,288-D space has room for vastly more than 12,288 almost-independent meanings.
If asked: “Precise statement?” For tolerance , you can fit on the order of unit vectors with all pairwise . That’s the Johnson–Lindenstrauss flavour of result. (3Blue1Brown’s video had a code bug in this demo; his lesson text includes the correction.)
MLP & facts · Superposition
Superposition: more features than neurons
Elhage et al., “Toy Models of Superposition” (2022); Bricken et al., “Towards Monosemanticity” (2023); Templeton et al., “Scaling Monosemanticity” (2024), all Anthropic
~1.5 min
A feature is a direction in activation space that means something: “is French”, “is code”, “mentions DNA”.
Tidy world: two neurons, two features, each feature on its own axis. One neuron = one meaning.
But the model wants far more features than it has dimensions. Trick: pack five directions into two dimensions. They overlap a little (cosine 0.31), which causes small interference.
Consequence: a single neuron responds to several unrelated features. That is polysemanticity , and it’s why reading neurons one by one fails.
Interpretability tool: a sparse autoencoder learns a much wider dictionary in which each activation uses only a few directions. Those directions are often clean, human-readable features.
Intuition: sparsity makes overlap cheap: if features rarely co-occur, they can share space.
Analogy: radio stations on nearby frequencies: fine as long as only a few broadcast at once.
If asked: “Does this work at scale?” Anthropic’s 2024 work extracted millions of features from a production model, including a famous Golden Gate Bridge feature that, when amplified, made the model obsessed with the bridge.
MLP & facts · How the pieces divide the work
Attention moves information, MLPs transform it
~1 min
Put it together on a fact-recall prompt: “The Eiffel Tower is in” → ?
Attention moves information between lanes: early on, “Eiffel” into “Tower” (build the entity); later, the entity into the last position.
MLPs work within a lane: at “Tower”, look up attributes of the Eiffel Tower (in Paris, a landmark).
Rough depth trend: early layers handle tokens, spelling and local syntax; middle layers build entities and recall facts; late layers decide what to say next.
The last lane now points toward “Paris”, and the unembedding (Part E) reads that off.
Intuition: attention = communication, MLP = computation.
If asked: “Is the depth story proven?” It’s a tendency seen across probing and interpretability studies (e.g., Geva et al. 2023 traced this exact subject → attribute flow), not a hard rule. Layers overlap a lot.
PART E
From the last vector to the next word
Unembedding, softmax, temperature, training, and the whole stack in one pass.
Putting it together · Output
Unembedding: the last vector scores every token
~1 min
After 96 layers, take the vector in the last position (“fox”).
This vector has been edited by every block; it should now point toward what comes next.
The unembedding matrix has one row per vocabulary token. Each row dotted with the vector gives that token’s score, a logit . 50,257 logits.
Softmax turns logits into probabilities. Negative logits become tiny probabilities, not negative ones.
Another 617M weights. Numbers here are illustrative.
Intuition: each vocabulary token has a direction; the next token is whichever direction the final vector points along most.
If asked: “Do other positions get unembedded?” During training, yes: every position predicts its next token (that’s the masking payoff). At generation time only the last one matters.
Putting it together · Sampling
Temperature controls how adventurous sampling is
: the model’s own distribution.
: sharper. Safer, more repetitive.
: always the top token (greedy).
: flatter. More surprising, more mistakes.
~1 min
Before sampling we can reshape the distribution with one knob: temperature . Divide every logit by , then softmax.
: exactly what the model learned.
: the gaps between logits are magnified, so the top choice dominates.
: all mass on the argmax. Deterministic, often dull and loopy for long text.
: gaps shrink, the tail gets real probability. Creative, then incoherent.
Intuition: temperature trades reliability for variety.
Analogy: physics: low temperature, the system sits in its lowest-energy state; high temperature, it explores. The name is borrowed from the Boltzmann distribution.
If asked: “Other knobs?” Top-k and top-p (nucleus) sampling cut off the tail before sampling. Same goal: keep variety without the garbage.
Putting it together · Training
Training is deck 2’s loop on next-token guesses
GPT-3: 300 billion training tokens. Llama 3: over 15 trillion .
Sources: Brown et al. 2020 (300B tokens); Meta Llama 3 announcement via The Register, Apr 2024 (15T+ tokens)
~1.2 min
Take real text. Thanks to masking, every position is a prediction problem: seven examples from one eight-word sentence.
Loss = negative log of the probability assigned to the true next token (cross-entropy). An untrained model is roughly uniform: , loss .
Then deck 2’s loop: backprop gives the gradient of this loss with respect to all 175B weights; take a step; repeat. The bars move as training proceeds. Some tokens stay hard (“verdant” is genuinely unpredictable).
Scale: GPT-3 saw 300B tokens; Llama 3 over 15 trillion. Same objective throughout.
Intuition: “predict the next token” sounds trivial, but doing it well on all of the internet forces the model to learn grammar, facts, and some reasoning.
If asked: “Plain SGD?” GPT-3 used Adam, a momentum-plus-per-weight-step-size variant of gradient descent, with large minibatches (3.2M tokens). Conceptually still deck 2.
Putting it together · End to end
The whole stack, end to end
~1.5 min
One full forward pass, everything from today.
Tokenize: four tokens, four IDs.
Embed: each ID becomes a 12,288-D vector (plus its position). Four lanes of the residual stream.
Layer 1: attention lets lanes exchange information (causally: each looks only back), then the MLP processes each lane. Both add deltas.
Layers 2 through 96: the same thing 95 more times, each with its own weights.
Unembed the last lane, softmax: a distribution with “ jumps” on top.
Sample, append, and run it all again for the next token.
Intuition: embed → (communicate, compute) × 96 → read off the next token.
If asked: “Do we recompute everything for each new token?” Naively yes. The KV cache stores past keys and values so only the new position is computed. Deck 4.
Putting it together · Recap
Six ideas to carry into deck 4
1 · Tokens
Text → integer IDs The model never sees letters.
2 · Embeddings
ID → vector Directions in 12,288-D carry meaning.
3 · Residual stream
One vector per position Every block reads it and adds an edit.
4 · Attention
Moves information Queries match keys; values get added. Many heads in parallel.
5 · MLPs
Process in place Two thirds of the weights; superposition packs in features.
6 · Output
Logits → softmax → sample Trained by gradient descent on next-token prediction.
~1 min
Tokens: integers. Strawberry is one symbol.
Embeddings: meaning lives in directions.
Residual stream: the backbone; everything is an additive edit.
Attention: communication between positions. .
MLPs: computation within a position; facts; superposition.
Output: unembed, softmax with temperature, sample. Training: next-token cross-entropy plus backprop.
Intuition: if someone can explain these six cards, they understand a GPT.
If asked: “What changed between GPT-3 and today’s models?” Mostly scale, data, better position encodings (RoPE), mixture-of-experts MLPs, and post-training. The skeleton on these cards is the same.
Putting it together · Next
Watch these, then let’s scale it up
3Blue1Brown, Deep Learning (Grant Sanderson)
Deck 4: serving, scale, RL and test-time compute.