QMIND Research · Deck 4 of 7
The frontier: serving, scale, RL
The same transformer from deck 3, now running for millions of people: what it costs to serve, what it takes to train, and how RL and extra thinking time taught it to reason.
A · Serving
B · Scale
C · Reinforcement learning
D · Test-time compute
~1 min
Deck 3 ended with a working transformer. This deck asks: what happens when you serve it to millions of people, and train it on a city's worth of GPUs?
Part A, serving: why generating text is memory-bound, the KV cache, and why attention is quadratic.
Part B, scale: GPUs, NVIDIA, the parallelism tricks for training, and the money and power being spent.
Part C, RL: how a next-token predictor becomes an assistant and then a reasoner. I'm learning this one too, so we'll build it from first principles.
Part D, test-time compute: getting better answers by spending more compute when answering. Then one slide on why all this pushed our project elsewhere.
Intuition: the frontier is now as much an engineering and capital problem as a modelling one.
If asked: "Is this on the mini-presentation list?" Yes: KV cache, parallelism, RLHF/GRPO and test-time compute all make good 5-minute talks.
PART A
Serving: from weights to a product
Every token a chatbot sends you is one more trip around a loop. The loop's bottleneck is memory, not math.
Serving · the generation loop
Reading is parallel; writing is one token at a time
Prefill
Whole prompt in one pass. Big matrix–matrix multiplies. Compute-bound. Sets time to first token .
Decode
One position per pass, hundreds of passes. Matrix–vector multiplies. Memory-bound. Sets tokens per second .
~1.5 min
Recall deck 3: the model outputs a probability distribution for the next token. To write a paragraph we call it in a loop.
Prefill: all six prompt tokens go through the network together. Each layer is one big matrix multiply over six rows, fully parallel. The last position's output gives the first new token.
Decode: now only the newest token goes in. The model picks from the distribution (temperature from deck 3), appends, and repeats.
This repeats once per output token. A 500-token answer means 500 sequential passes. You cannot parallelise across them: token 7 depends on token 6.
So serving has two phases with opposite bottlenecks. Prefill keeps the math units busy. Decode mostly waits on memory, which is the next slide.
Intuition: reading a prompt is a batch job; writing an answer is a relay race with one runner at a time.
Subtle point: during decode the new token still attends to all earlier tokens. It just doesn't recompute them, thanks to the KV cache (coming up).
If asked: "Why does the first token take longer?" Prefill processes the whole prompt; for a long document it's a lot of FLOPs. After that, each token costs about the same.
Serving · the memory wall
Decode is memory-bound: each token reads every weight
Llama‑3.1‑8B in BF16 is 16 GB of weights. Each decode step streams all of it from HBM to the math units.
That caps one user at about 210 tokens/s , however fast the chip computes.
The math needed: GFLOP, which takes 16 µs at ~989 TFLOP/s. Tensor cores are busy about 0.3% of the time.
So read the weights once and use them for many users at the same time.
H100 SXM: 80 GB HBM3 at 3.35 TB/s, ~989 dense BF16 TFLOP/s (NVIDIA spec sheet). Llama 3.1 8B ≈ 8.0B params. Upper-bound arithmetic, as of Oct 2026.
~1.5 min
The picture: weights live in HBM (high-bandwidth memory, the GPU's main memory). The math happens in the tensor cores. A pipe connects them.
To produce one token, every weight matrix must be multiplied by the new token's vector. So every byte of weights travels through the pipe once per token.
Divide bytes by bandwidth: 16 GB / 3.35 TB/s ≈ 4.8 ms. That's a hard floor for batch size 1, about 210 tokens/s.
The arithmetic is tiny: about 2 FLOPs per parameter per token (one multiply, one add), so 16 GFLOP, which is 16 µs of compute. The chip spends over 99% of the time waiting for data.
The fix is batching: one trip of the weights through the pipe can be multiplied against many users' vectors at once.
Intuition: the chef (tensor cores) is incredibly fast but the ingredients (weights) arrive on one conveyor belt. Cooking for one diner wastes the chef.
Common mistake: people quote FLOPs to estimate chat speed. For single-user decode, bandwidth is what matters. That's why HBM bandwidth is the headline spec of each new GPU.
If asked: "Do caches help?" On-chip SRAM is ~tens of MB, far smaller than 16 GB of weights, so weights can't stay on chip between tokens.
Serving · the KV cache
KV cache: compute keys and values once, keep them
A new token's query must meet the keys and values of every earlier token, in every layer.
Recomputing them at every step repeats work we already did.
So store them. Prefill writes a column for each prompt token.
Each decode step adds one new column : its K and V for every layer and KV head.
We traded compute for memory. The cache grows with every token.
~1.5 min
Deck 3 recap: attention for a token computes with every earlier key, softmaxes, and mixes the earlier values. The amber query is the newest token looking back.
Without a cache, each new step would push the whole sequence through again, recomputing K and V for every past token in every layer. Rose = wasted work, growing every step.
The key observation: because of causal masking, token 3's K and V never change once computed. Later tokens can't affect them. So we cache them.
Each decode step computes K and V for the new token only (one column) and appends it. Then the new query attends over the whole cached block.
Cost: memory. The cache holds 2 (K and V) × layers × KV heads × head dimension numbers per token. Next slide: how big that gets.
Intuition: it's memoisation. Past tokens are frozen, so their intermediate results can be stored and reused.
Subtle point: only K and V are cached, not Q. Past queries are never needed again: only the newest token is asking the question.
If asked: "Why does causality matter?" In a bidirectional model (BERT), a new token would change every earlier representation, so you couldn't cache. Causal masking is what makes the KV cache valid.
Serving · cache arithmetic
How big is the cache? Do the arithmetic
Llama‑3.1‑70B: 80 layers, 8 KV heads, , BF16:
32 users at 8K tokens each: 80 GiB of cache, more than a whole H100.
Without grouped KV heads (64 instead of 8): 8× larger, 2.5 MiB per token.
Config: Llama‑3.1‑70B (80 layers, 64 query / 8 KV heads, head dim 128), per Meta's released config; KV sizing as in MinIO MemKV docs. As of Oct 2026.
~1.5 min
Let's size it for a real open model.
The 2 is for K and V. Then one vector of size per layer, per KV head, per token, times bytes per number.
Llama‑3.1‑70B: bytes ≈ 320 KiB for every single token of context.
At its full 128K context that's about 40 GiB for one conversation. Compare with the weights: 70.6B params × 2 bytes ≈ 141 GB, which already needs two H100s.
Serving many users multiplies it: 32 users × 8K tokens ≈ 80 GiB. The KV cache, not the weights, is often what limits how many users fit on a GPU.
Llama uses grouped-query attention: 64 query heads share 8 KV heads. With full multi-head attention the cache would be 8× bigger. We'll see this trick in a few slides.
Intuition: the weights are a fixed cost shared by everyone; the KV cache is a per-user, per-token cost that grows as conversations get longer.
Watch the units: KiB/GiB are powers of 2 (1 GiB ≈ 1.074 GB). 40 GiB ≈ 43 GB.
If asked: "Why do long-context APIs charge more?" Long contexts eat KV memory, so fewer users fit per GPU, and attention over long contexts costs more compute.
Serving · why long context is expensive
Attention is quadratic in context length
One pass over a whole sequence: work for the scores and the mixing.
With the cache, one new token costs : one new row .
Generating tokens is still quadratic overall.
Memory only grows linearly: the cache is one row of entries per layer and head.
~1.5 min
The grid is the attention score matrix: row i = token i's query, column j = token j's key. Causal mask means only the lower triangle is used.
For tokens, is . Double the context, four times the scores.
Each score is a d-dimensional dot product, and mixing values costs the same again, so a full pass is .
During decode we only compute the newest row (amber): the new query against all T cached keys. That's linear in T per token.
But step t costs t, so the total is : the area of the triangle. Long generations (like reasoning traces) are quadratic.
Memory is the bar underneath: the cache stores T keys/values, linear. We never store the full matrix during decode.
Intuition: every new guest at a party has to say hello to everyone already there. Each arrival is linear; the whole party is quadratic.
Subtle point: per token, the weight matmuls cost about FLOPs regardless of , while attention costs about . For Llama‑70B these are equal near tokens. Below that the MLPs dominate FLOPs; above it attention does. In decode, reading the KV cache can dominate memory traffic much earlier once you batch many users.
If asked: "So why do 1M-token contexts exist?" Tricks on the next slides (FlashAttention, sparse/sliding attention, hybrids) plus a lot of hardware.
Serving · batching and cost
Batching: one weight read serves many users
Stack B users' vectors into a matrix: one trip of the weights through memory now produces B tokens .
Total throughput climbs almost linearly at first…
…while each user's speed drops slowly.
Cost per million tokens falls about 20× from B = 1 to B = 64.
Then the KV caches fill memory, and reading them dominates. That is why cache size matters so much.
Illustrative roofline model, ideal kernels: Llama‑3.1‑8B, BF16, one H100, 4K‑token contexts, $2.50/GPU‑hr (H100 rentals ≈ $1.49–$6.98/hr, IntuitionLabs 2026).
~1.5 min
Matrix–vector becomes matrix–matrix. The weights still travel once per step, but now they meet B vectors. This is the single most important serving trick.
Blue: total tokens/s across all users. At small B, step time barely changes (still waiting on weights), so throughput ≈ B × 200 tokens/s.
Amber: tokens/s each user sees. It falls because each step now also reads every user's KV cache.
Green: dollars per million output tokens = GPU price ÷ tokens per hour. Batching is why API tokens are cheap: ~$3.4/M at B=1 vs ~$0.16/M at B=64 in this toy model.
The wall: at 4K tokens each user's cache is ~0.5 GB; about 119 users fill the H100. And past ~30 users, KV reads exceed weight reads, so throughput flattens.
Intuition: a bus costs about the same to drive with 1 passenger or 60. Batching fills the bus; the KV cache is each passenger's luggage.
Honesty note: these curves come from a simple formula (step time = bytes moved ÷ bandwidth, or FLOPs ÷ peak, whichever is larger), not a benchmark. Real servers get maybe 50–80% of these numbers.
If asked: "Why are output tokens priced higher than input tokens?" Input tokens are processed in prefill (parallel, compute-efficient); output tokens each need a memory-bound decode step.
Serving · scheduling and memory
Continuous batching and paged KV memory
vLLM (2023): earlier servers wasted 60–80% of KV memory to fragmentation and over‑reservation; paging cut waste to under 4% , for up to 24× the throughput of HF Transformers.
Source: vLLM project blog, "Easy, Fast, and Cheap LLM Serving with PagedAttention", Jun 2023 (Kwon et al., SOSP 2023).
~1.5 min
Top: four batch slots over time. Requests have different lengths. With static batching, a batch runs until its longest request finishes.
Grey = idle slots. A short answer finishes and its slot sits empty, waiting for the long one.
Continuous (in-flight) batching: the scheduler checks after every decode step and slots a waiting request in as soon as one finishes. Same work, done sooner, GPU kept full.
Bottom: memory. Naively each request reserves a contiguous chunk big enough for its maximum length. Most of it sits empty (rose), and gaps between chunks fragment memory.
PagedAttention borrows the operating-system idea of virtual memory: split the cache into fixed blocks (e.g. 16 tokens), allocate on demand, anywhere in memory, and keep a block table per request, like a page table.
Result: almost no waste, so more requests fit, so bigger batches, so more throughput. Same model, same GPU.
Intuition: don't book a 10-seat table for every party "just in case"; seat people as they arrive and give each party chairs wherever there's room.
Bonus: paging also lets requests that share a prefix (same system prompt) share the same physical blocks. That is "prefix caching".
If asked: "Is vLLM what companies use?" vLLM and SGLang are the main open-source servers; labs run their own stacks built on the same ideas.
Serving · bending the curve (1)
Shrink the bytes: share KV heads, use fewer bits
GQA: groups of query heads share one K/V head. Llama‑3.1‑70B: 64 query heads, 8 KV heads, so the cache is 8× smaller.
MQA shares a single K/V head. MLA (DeepSeek) compresses K and V into a small latent vector.
Format Bytes 70B weights
BF16 2 141 GB
FP8 / INT8 1 71 GB
FP4 / INT4 0.5 35 GB
Fewer bytes to stream means faster decode and more users per GPU. The price is some accuracy, so it needs care.
GQA: Ainslie et al. 2023; MQA: Shazeer 2019; MLA: DeepSeek‑V2 2024 (reports a 93.3% smaller KV cache than DeepSeek 67B). Native FP8 on Hopper, FP4 on Blackwell.
~1.5 min
Standard multi-head attention: every query head has its own K and V head. The diagram shows 8 heads standing in for Llama's 64.
Grouped-query attention: several query heads share one K/V head. Each head still asks its own question but reads from a shared memory. The cache shrinks by the group size, and quality barely moves.
Multi-query is the extreme case, one shared K/V. Multi-head latent attention (DeepSeek‑V2/V3) caches a compressed latent and expands it on the fly.
The other lever is bytes per number. Weights in FP8 halve the memory traffic; 4-bit halves it again. KV caches can be quantised too.
Because decode is memory-bound, halving bytes roughly doubles single-user speed. Newer GPUs have hardware for low-precision formats (FP8 on H100, FP4 on Blackwell).
Intuition: everything in Part A comes back to "how many bytes cross the bus per token". Each trick is a way to move fewer.
Common mistake: quantisation isn't free. 4-bit weights can hurt reasoning and long-context tasks. Teams measure before shipping.
If asked: "Why does GQA barely hurt quality?" Heads turned out to have a lot of redundancy in their keys/values. Many queries can share them, and models are trained (or uptrained) with GQA from the start.
Serving · bending the curve (2)
Fewer memory trips, or less attention
Fixed-size memory forgets. Many recent models are hybrids : mostly cheap layers, plus a few full-attention layers.
FlashAttention: Dao et al. 2022 (exact, IO-aware). Sliding window: e.g. Mistral 7B (2023). State-space: Mamba, Gu & Dao 2023.
~1.5 min
Left: standard attention writes the whole score matrix out to HBM, reads it back for softmax, writes it again. For long T, that traffic dominates.
FlashAttention: load a tile of Q and a tile of K,V into fast on-chip SRAM, compute that block of scores, update a running softmax, accumulate the output, discard the tile. The matrix never exists in HBM.
Same math, exact same result, same FLOPs. But memory drops to and HBM traffic falls a lot, so it's several times faster in practice.
Right: other approaches change the math. Full causal attention: every token sees every earlier token.
Sliding window: each token only sees the last w tokens. Cost , and the cache is capped at .
Linear attention and state-space models (Mamba) keep a fixed-size state updated per token, like an RNN. O(T) time, O(1) memory per sequence.
The catch: a fixed-size state can't remember everything exactly, which hurts tasks like recalling a phone number from page 3. Hybrids keep a few full-attention layers for that.
Intuition: FlashAttention is a better kitchen layout (same recipe, fewer trips to the fridge). Windows and SSMs are a different recipe.
Common mistake: "FlashAttention makes attention linear." No: compute is still quadratic. It removes the quadratic memory and most of the slow memory traffic.
If asked: "What's the 'online softmax'?" Keep a running max and running sum per row; when a new tile has a bigger max, rescale what you've accumulated. That's what lets softmax be computed tile by tile.
Serving · bending the curve (3)
Speculative decoding: draft cheaply, verify in one pass
Verification uses a rejection-sampling rule, so the output distribution is exactly the big model's. Reported speedups: about 2–3× .
Leviathan, Kalman & Matias, "Fast Inference from Transformers via Speculative Decoding", ICML 2023; Chen et al. (DeepMind) 2023.
~1.25 min
Decode is memory-bound, so the big model has spare compute. Checking k tokens in one pass costs about the same as generating one, just like prefill.
A small draft model (fast because it's small) guesses the next 4 tokens one by one.
The big model runs once over all 4 guessed positions in parallel and gets its own distribution at each position.
Accept guesses left to right while they pass the test; at the first rejection, sample a replacement from a corrected distribution and throw away the rest. Here we keep 3 and fix the 4th.
So 4 tokens for one big-model pass instead of four. Easy text ("of the", code boilerplate) gets accepted often; hard tokens fall back to normal speed.
Intuition: a junior writes a draft, the senior reads it once and signs off on everything up to the first mistake. The senior reading is much cheaper than the senior writing.
Subtle point: the rule accepts a draft token with probability and on rejection resamples from the normalised . That makes the final samples distributed exactly as the big model would produce. It's not an approximation.
If asked: "Do you need a separate model?" Not always: variants like Medusa or EAGLE add small extra heads to the big model itself to make the draft.
PART B
Scale: chips, wires and gigawatts
Why GPUs, why NVIDIA, how one model is split across tens of thousands of chips, and what the build-out costs.
~0.25 min
Part A was one model on one or a few GPUs. Part B: what changes when training needs 16,000 GPUs and serving needs millions.
Intuition: at frontier scale the hard problems are moving data between chips and getting enough electricity.
Scale · why GPUs
Neural nets are matrix multiplies; GPUs are built for them
A layer is : millions of independent multiply–adds.
A CPU has a few big cores built for branchy, serial code. It walks through the output a few tiles at a time.
A GPU runs thousands of simple lanes doing the same operation on different data. H100: 132 SMs, 528 tensor cores.
A tensor core does a whole small matrix multiply–accumulate per instruction: .
Transformer FLOPs are overwhelmingly matmuls, so the fit is near perfect.
H100 SXM5: 132 streaming multiprocessors, 4 tensor cores each (NVIDIA Hopper architecture whitepaper, 2022).
~1.25 min
Remember the MNIST network from deck 2: every layer was weights times activations. Transformers are the same: QKV projections, attention, MLPs are all matrix multiplies.
Each output entry of C = AB is an independent dot product. A CPU has maybe 8–64 cores and spends its transistors on branch prediction and big caches for serial code.
A GPU spends its transistors on arithmetic: many simple cores in lockstep (SIMT). Every output tile can be computed at once.
Tensor cores go further: one instruction multiplies small matrices (e.g. 16×16 tiles) and accumulates. That's where the hundreds of TFLOP/s come from, in low precision (BF16/FP8/FP4).
Matmuls also have high arithmetic intensity: an matmul does about FLOPs on numbers, so there's lots of reuse per byte loaded. That matters on the roofline slide.
Intuition: a CPU is a few brilliant generalists; a GPU is a stadium of people who can each do one multiply-add, all at once, on command.
If asked: "Why did GPUs exist before AI?" Graphics is also massively parallel: millions of pixels, each shaded independently. Deep learning found a matching workload around 2012 (AlexNet trained on two consumer GPUs).
Scale · why NVIDIA
NVIDIA's moat is software and systems, not just chips
Software · since 2007
CUDA and its libraries cuBLAS, cuDNN, NCCL, TensorRT. PyTorch runs best on CUDA. Two decades of tuned kernels.
Interconnect
NVLink, NVSwitch, InfiniBand GPU-to-GPU wires inside the rack, plus Mellanox networking (acquired 2020) between racks.
Systems
Sell the rack, not the chip GB200 NVL72: 72 GPUs in one NVLink domain that behaves like one giant GPU.
HBM per GPU: H100 3.35 TB/s; H200 4.8; B200 8; GB300 8; Rubin ~22 (HBM4, detailed Jan 2026, shipping H2 2026). NVIDIA specs/announcements, as of Oct 2026.
~1.5 min
NVIDIA isn't the only one making fast chips: Google has TPUs, AMD has Instinct GPUs, Amazon has Trainium. So why does NVIDIA dominate?
Software: CUDA launched in 2007. Every framework, paper and kernel was written for it first. Switching means re-tuning your whole stack. That's a moat built from developer time.
Interconnect: training splits a model across GPUs, so the wires between them matter as much as the chips (coming up). NVIDIA owns NVLink and bought Mellanox for InfiniBand.
Systems: they ship whole racks. NVL72 puts 72 GPUs on one NVLink fabric, so model parallelism that used to stop at 8 GPUs now spans 72.
Right: memory bandwidth per GPU by generation. Given Part A, this is the number that sets decode speed. It has grown ~2.4× from H100 to B200.
Rubin (detailed at CES 2026, shipping second half of 2026) moves to HBM4 at a claimed ~22 TB/s and 288 GB per GPU.
Intuition: NVIDIA sells a solution (chips, wires, software, racks) that already works at 100,000-GPU scale. Competitors have to match all of it.
Caveat: Rubin numbers are NVIDIA's announced figures; independent measurements come once systems are deployed. Capacity: 80 / 141 / ~180–192 / 288 / 288 GB.
If asked: "Will the moat last?" Google trains Gemini on TPUs, and labs are funding custom chips, so it's being tested. But CUDA plus rack-scale systems is still the default in 2026.
Scale · the roofline model
Arithmetic intensity decides the bottleneck
A weight matmul at batch : FLOPs over bytes, so .
Above the ridge (~295 FLOP/byte on H100) you are compute-bound . Below it, memory-bound .
Attention over each user's KV cache stays near at any batch size: every user brings their own cache.
Roofline model: Williams, Waterman & Patterson, CACM 2009. H100 SXM figures: 3.35 TB/s, ~989 dense BF16 TFLOP/s; ridge = 989 / 3.35 ≈ 295 FLOP/byte.
~1.5 min
Axes, both log scale: x = arithmetic intensity (FLOPs per byte fetched from memory), y = achievable FLOP/s.
The slanted line: if each byte supports only I FLOPs, you can't go faster than . That's the memory-bound region.
The flat roof: you can't exceed peak FLOPs. They meet at the ridge point, ~295 FLOP/byte for an H100.
Decode at batch 1: intensity ≈ 1, so we reach about 3 TFLOP/s out of 989. That's the 0.3% from earlier, seen on a chart.
Batch 32: intensity ≈ 32, so 32 times further up the slope. Still memory-bound, but much better.
Prefill of a long prompt, or a big training batch: intensity in the hundreds to thousands, so we hit the roof.
The catch from the batching slide: attention reads a separate KV cache per user, so batching doesn't raise its intensity. That's why KV-cache size dominates large-batch serving.
Intuition: intensity measures how much cooking you do per ingredient delivered. Low intensity: you wait for deliveries. High intensity: you're limited by how fast you cook.
If asked: "Where does SRAM fit?" On-chip SRAM (~50 MB L2 plus 228 KB shared memory per SM on H100) is much faster than HBM. FlashAttention and tiled matmuls raise effective intensity by reusing data while it's on chip.
Scale · training
A frontier model won't fit on one GPU, or finish on one
Mixed-precision Adam keeps about 16 bytes per parameter : weights, gradients, FP32 master copy, two moments.
70B parameters → 1.13 TB of state, before activations: about 15 H100s just to hold it.
Llama 3.1 405B used FLOPs.
Time at 40% utilisation
3,000 years
16 B/param: ZeRO paper (Rajbhandari et al. 2020). FLOPs and up to 16,384 H100s at 38–43% MFU: "The Llama 3 Herd of Models", Meta, Jul 2024.
~1.25 min
Two separate walls: memory and time.
Training state per parameter: 2 bytes BF16 weights + 2 bytes gradients…
…+ 4 bytes FP32 master weights (small updates get lost in BF16)…
…+ 4 + 4 bytes for Adam's first and second moments. Total 16 bytes.
A 70B model needs ~1.13 TB, against 80 GB per H100. And activations for backprop come on top.
Time: training FLOPs , params × tokens × 6 (2 for the forward pass, 4 for backward). Meta reports for the 405B model.
One H100 at a realistic 40% of peak would take about 3,000 years. 16,384 of them: about two months, if nothing breaks. So we must split the work. Next four slides: how.
Intuition: the model is too big to fit in one place and the job too long to run in one place, so both data and model have to be split.
Where 6ND comes from: forward ≈ 2 FLOPs per param per token (multiply + add). Backward ≈ twice forward: gradients w.r.t. activations and w.r.t. weights.
If asked: "What's MFU?" Model FLOPs utilisation: useful FLOPs ÷ peak. 40% is good at this scale, because communication, memory stalls and failures eat the rest.
Scale · data parallelism
Data parallel: copy the model, split the batch, average
Ring all-reduce: each GPU sends about × the gradient size, nearly independent of .
~1.5 min
Four GPUs, each with a full copy of the model. This is the simplest and most common form of parallelism.
Split the minibatch: each GPU gets a different quarter of the examples (blue packets).
Each GPU runs forward and backward on its shard and gets its own gradient (amber, different heights because the data differ).
All-reduce: GPUs pass pieces of their gradients around a ring, adding as they go, until every GPU holds the same average. That average equals the gradient of the full batch: SGD from deck 2, just computed in pieces.
Every GPU applies the same update to the same weights, so the copies stay identical without ever sending weights.
Cost: each GPU sends and receives about twice the gradient size per step, regardless of how many GPUs. And it can overlap with the backward pass (send layer L's grads while computing layer L−1).
Intuition: four students each mark a quarter of the exams, then pool their tallies so everyone has the class average.
How the ring works: split the gradient into N chunks. In N−1 "reduce-scatter" steps each GPU ends up owning the full sum of one chunk; in N−1 "all-gather" steps those sums are passed around. Each step sends 1/N of the data.
If asked: "What's the limit?" Every GPU still holds the full model and optimizer state. That's the next slide.
Scale · sharding the state
ZeRO / FSDP: shard the state instead of copying it
A 7B model with Adam needs ~112 GB per copy. Plain data parallel copies it to every GPU.
Stage 1: each GPU keeps only its 1/N slice of the optimizer state.
Stage 2: gradients too: reduce-scatter instead of all-reduce.
Stage 3 / FSDP: parameters too. Each layer is gathered just in time, used, then freed.
The price is more communication. It fits because it overlaps with compute.
Rajbhandari et al., "ZeRO: Memory Optimizations Toward Training Trillion Parameter Models", 2020; PyTorch FSDP (Zhao et al. 2023).
~1.25 min
Four GPUs' memory, with the 80 GB line. Blue = BF16 params (14 GB), amber = grads (14 GB), violet = optimizer state (84 GB). Plain data parallel: doesn't fit.
The optimizer state is only needed during the update, and each GPU can update just its own quarter of the parameters. So shard it: 84 → 21 GB each. Now it fits.
Gradients: instead of every GPU getting the full averaged gradient, use reduce-scatter so each gets only the slice it will update.
Parameters: store only a slice; before running layer k, all-gather its full weights, compute, throw them away. Memory per GPU ≈ 112 GB / plus one layer at a time.
Stage 3 adds an all-gather in forward and backward, about 1.5× plain data-parallel traffic. With fast links and prefetching it hides behind compute.
Intuition: instead of every student carrying the whole textbook, each carries a few chapters and they pass pages around when needed.
Naming: ZeRO is DeepSpeed's (Microsoft) name; FSDP (Fully Sharded Data Parallel) is PyTorch's equivalent of stage 3. Deck 7 will touch FSDP in PyTorch.
If asked: "Does it change the maths?" No. It's still exactly data-parallel SGD/Adam; only where the bytes live changes.
Scale · tensor parallelism
Tensor parallel: split each matrix across GPUs
An all-reduce in every layer, forward and backward. It only pays inside a fast NVLink domain.
~1.25 min
When even one layer's matrices are too big or slow for one GPU, split the matrices themselves. This is Megatron-LM style tensor parallelism.
First matmul (e.g. the MLP's up-projection): split W by columns. Each GPU computes its slice of the output, with no communication, because each column of the output only needs one column of W.
The nonlinearity (GELU) is elementwise, so each GPU applies it to its own slice. Still no communication.
Second matmul (down-projection): split W′ by rows to match. Each GPU produces a partial sum of the full output, so we need an all-reduce to add them up.
One all-reduce per block in forward, one in backward, in every layer, on the critical path. That's why tensor parallelism stays within a node (8 GPUs) or an NVL72 rack.
Intuition: four people each multiply by a quarter of the columns, then add up their partial answers.
Attention: split by heads: each GPU owns some heads, which are independent, then the output projection is row-split with one all-reduce.
If asked: "Why column then row?" That ordering lets the elementwise nonlinearity happen locally, so the MLP needs one all-reduce instead of two.
Scale · pipeline parallelism
Pipeline parallel: layers across GPUs, and the bubble
stages, microbatches. Smarter schedules (1F1B, interleaved) also cut memory and bubbles.
~1.5 min
Split the model by depth: GPU 0 has the first quarter of the layers, GPU 1 the next, and so on. Only activations cross between stages, once per microbatch, so this works across slower links between nodes.
The problem: with one batch, GPU 1 waits for GPU 0's forward, GPU 0 waits for everyone's backward. Grey = idle. With 4 stages, 75% of the time is bubble.
Fix: cut the batch into microbatches and stream them through like an assembly line. GPU 0 starts microbatch 2 while GPU 1 works on 1. Bubble fraction .
More microbatches: with m = 16 the bubble drops to about 16% and total time shrinks toward the ideal.
Real systems use schedules like 1F1B (one forward, one backward alternating) so you don't hold all microbatches' activations at once, and interleaved stages to shrink the bubble further.
Intuition: an assembly line is efficient only when it's full. Startup and drain are the bubble; many small items amortise it.
Simplification: the chart draws backward the same length as forward. In reality backward takes about twice as long.
If asked: "Why not huge m?" Each microbatch must still be big enough to keep a GPU efficient, and the global batch size is limited by optimisation (too big a batch hurts sample efficiency).
Scale · mixture of experts
Mixture of experts: each token visits a few experts
The MLP becomes many experts . A small router scores them per token and keeps the top .
Experts live on different GPUs (expert parallelism ), so routing is an all-to-all shuffle.
DeepSeek‑V3: 671B parameters in total, 37B active per token (8 of 256 routed experts, plus 1 shared).
A balancing term keeps load even, so no expert or GPU becomes the bottleneck.
DeepSeek‑AI, "DeepSeek‑V3 Technical Report", Dec 2024. MoE layers: Shazeer et al. 2017; Switch Transformer, Fedus et al. 2021.
~1.5 min
Recall deck 3: the MLP block stores a lot of the "facts". MoE replaces one big MLP with many smaller ones and lets each token use only a few.
The router is a tiny linear layer plus softmax: it scores every expert for this token. Keep the top-k (here 2).
Output = gate-weighted sum of the chosen experts' outputs. Gradient flows to the chosen experts and through the gate scores, so the router learns too.
Experts sit on different GPUs, so each layer sends tokens to wherever their experts live and back again: two all-to-all exchanges per MoE layer.
The payoff: parameters (knowledge capacity) and FLOPs per token are decoupled. DeepSeek‑V3 stores 671B params but each token pays for about 37B.
The failure mode: the router sends everything to a few favourite experts. Auxiliary balancing losses (or DeepSeek's bias adjustment) keep the load even.
Intuition: a hospital with many specialists. Triage sends each patient to two of them; nobody sees every doctor.
Serving catch: all experts must sit in memory even though each token uses a few, so MoE saves FLOPs, not memory.
If asked: "Do experts specialise in topics?" Partly. Analyses find specialisation by token type and syntax as much as by topic. They're not neat "maths expert" / "French expert" units.
Scale · putting it together
Real runs combine all of these, shaped by the wires
Llama 3.1 405B: 16,384 H100s = TP 8 × PP 16 × DP 128.
HBM (H100) 3.35 TB/s
NVLink 5 / GPU 1.8 TB/s
NVLink 4 / GPU 0.9 TB/s
InfiniBand NDR / NIC 0.05 TB/s
Bars on a log scale.
Rule: put the chattiest parallelism on the fastest wire.
Llama 3 Herd of Models (Meta, Jul 2024), Table 4. NVLink 4 (H100) 900 GB/s; NVLink 5 (Blackwell) 1.8 TB/s (NVLink: bidirectional totals); NDR InfiniBand 400 Gb/s per port.
~1.25 min
A toy cluster: 16 nodes of 8 GPUs. Each square is a GPU.
Tensor parallel inside each node (violet outline): the all-reduce-every-layer traffic stays on NVLink.
Pipeline stages across nodes (colours): only activations cross, once per microbatch, so slower network links are fine.
Data parallel across whole pipeline replicas (blue outlines): one gradient sync per step, overlapped with backward.
Meta's real config for Llama 3.1 405B: TP 8 × PP 16 × DP 128 = 16,384 H100s, plus context parallelism (splitting the sequence) for 128K-token training.
Why that layout: each step down the bandwidth ladder is a big drop. Memory > NVLink > network: about two orders of magnitude end to end.
So the design rule: the more often a kind of parallelism talks, the faster the wire it needs. Rack-scale NVLink (NVL72) widens the fast domain from 8 GPUs to 72.
Intuition: organise a company so people who talk every minute share an office, and people who talk once a week can be in different cities.
"4D": data, tensor, pipeline, plus context/sequence parallel or expert parallel. People say 3D, 4D or 5D depending on what they count.
If asked: "What breaks at this scale?" Hardware. Meta reports frequent job interruptions over a 54-day window, mostly GPU and memory faults, so checkpointing and fast restart are part of the design.
Scale · why anyone spends this
Scaling laws: loss falls predictably with compute
The best model for each budget lies on a smooth line: a power law in compute.
Chinchilla: grow parameters and tokens together, about 20 tokens per parameter . 70B on 1.4T tokens beat 280B Gopher.
Today's models train far past that (Llama 3 8B: 15T tokens) because a smaller model is cheaper to serve.
Curves use Chinchilla's fitted constants (E=1.69, A=406.4, B=410.7, α=0.34, β=0.28; Hoffmann et al. 2022). Kaplan et al. 2020. Llama 3: Meta, 2024.
~1.5 min
Axes: training compute (log) vs pretraining loss. This is the empirical fact that justifies everything on the next three slides.
Each curve is one model size trained on more and more tokens. Small models improve, then flatten: they run out of capacity.
Bigger models start worse at small compute but keep going.
The lower envelope (best model for each budget) is close to a straight line on log-log axes: a power law. Ten times the compute buys a predictable drop in loss.
Chinchilla (DeepMind, 2022) fit the formula on the right and found earlier models were too big and undertrained: for a fixed budget, scale params and tokens roughly equally, about 20 tokens per parameter.
Since then labs deliberately "over-train" small models far past 20 tokens/param, because you pay training once and inference forever.
Intuition: scaling laws turn AI progress into an engineering forecast: "spend 10× more, get this much better". That predictability is what lets CFOs sign multi-billion-dollar cheques.
Careful: lower loss is not the same as specific capabilities. Some abilities look like they appear suddenly on benchmarks even while loss falls smoothly, and the law says nothing about which skills arrive when.
If asked: "Where does the 20 come from?" Minimising this fitted L subject to C = 6ND gives an optimal tokens/params ratio; the Chinchilla paper's three estimation methods land roughly at 20.
Scale · the build-out
The build-out, in dollars
Capital expenditure (mostly data centres, chips, power) of Amazon, Alphabet, Microsoft and Meta.
2026 guidance after Q2 earnings: about $720–745B , roughly 75% above 2025.
Amazon, Alphabet and Meta each raised their 2026 guidance during the year.
Plus Oracle (FY27 up to ~$95B) and the OpenAI–Oracle–SoftBank Stargate venture ($500B, >9 GW planned).
2024–25: company filings via Platformonomics (Feb 2026; Alphabet excl. finance leases). 2026: Q2 2026 earnings guidance (Jul 2026), per MLQ/IG summaries (Aug 2026). Oracle: Jun 2026.
~1.5 min
Capex = money spent on long-lived physical assets. For these companies it's now dominated by AI data centres.
2024: about $251B combined for the big four.
2025: about $416B. Amazon ~$135B, Microsoft ~$118B, Alphabet ~$91B, Meta ~$72B.
2026 guidance as of the July earnings: roughly $720–745B depending on how you add the ranges. Microsoft's ~$175B is after a lease-accounting change (it was ~$190B before; underlying plans unchanged), so I show a range.
Per company: Amazon ~$220B; Alphabet $195–205B; Microsoft ~$175B (~$190B before the lease change); Meta $130–145B. Amazon said it still can't meet AI demand; memory prices are also pushing costs up.
Then add Oracle (spent ~$56B in FY26, guiding up to ~$95B for FY27, part customer-funded), Stargate, xAI, CoreWeave, sovereign projects.
Intuition: four companies are each spending, per year, more than the GDP of many countries, almost all on one technology.
Definitions matter: some figures include finance leases and some don't; fiscal years differ (Microsoft's ends in June, Oracle's in May). Treat totals as ±5%.
If asked: "Is it all AI?" Not 100%: some is ordinary cloud, warehouses (Amazon), offices. But the companies say the growth is overwhelmingly AI.
Scale · power
The new unit of compute is the gigawatt
One gigawatt is about the output of a large nuclear reactor, running continuously.
World (IEA): data centres used ~415 TWh in 2024 (~1.5% of electricity), heading to ~945 TWh by 2030, a bit more than Japan uses today.
US (LBNL/DOE): 4.4% of electricity in 2023; 6.7–12% projected for 2028.
Power, grid connections and turbines, not chips, are now the slowest part to build.
Epoch AI, Stargate sites (Apr 2026); TechCrunch on Meta Prometheus/Hyperion (Jul 2025); IEA "Energy and AI" (2025); DOE/LBNL report (Dec 2024).
~1.25 min
Circle area = power. Labs now describe sites by gigawatts, not GPU counts.
Stargate's Abilene, Texas campus: about 0.3 GW operating as of April 2026 (~250k H100-equivalents), projected to reach ~1.2 GW by the end of 2026.
Meta: Prometheus (Ohio) ~1 GW coming online in 2026; Hyperion (Louisiana) planned to scale up to 5 GW over several years.
Stargate overall: over 9 GW planned across seven sites. Epoch compares that to New York City's peak demand.
Globally, the IEA projects data-centre electricity roughly doubling from 2024 to 2030.
In the US the share could go from ~4% to between ~7% and 12% within five years. That's a big jump for a grid that grew very little for 15 years.
So the binding constraints are now interconnection queues, transformers, gas turbines and permits. That's why labs are signing nuclear and gas deals.
Intuition: a GPU is a space heater that does maths. A million of them need a power plant and a cooling system to match.
Caveat: projections have wide ranges (note LBNL's 6.7–12%); planned gigawatts often slip.
If asked: "How much does one GW cost?" Estimates vary widely, often tens of billions of dollars per GW once GPUs are included. I'd quote a source before giving a number.
Scale · perspective
"Biggest build-out ever"? Depends on the yardstick
In nominal dollars: the largest private capital-spending wave on record.
As a share of GDP: about the same as the 1850s railroad boom. Wartime mobilisation was far larger.
What's different: chips wear out in years, while rails lasted a century. Revenue is still catching up, and more of it is debt-financed.
WSJ analysis via Benzinga, Feb 2026 (annual spend as % of GDP; AI = big-4 2026 est. $670B). Brookings (van Nieuwerburgh), Sep 2026: 3.6%/yr projected, 2025–32.
~1.5 min
Matt's framing is "biggest infrastructure build-out in history". Let's be precise about what's true.
In raw dollars, yes: no private sector has ever spent this much per year on one category of equipment.
But history should be compared as a share of the economy. The WSJ's comparison: Apollo ~0.2% of GDP per year, the interstate highways ~0.4%, 1850s railroads ~2%.
The big four's 2026 capex (estimated at $670B in February, since revised upward) ≈ 2.1% of US GDP: comparable to the railroads, below the Louisiana Purchase (~3%, a one-off). A Brookings projection that counts the whole build-out to 2032 gets ~3.6%/yr. WWII-era defence spending was an order of magnitude larger.
The differences cut both ways: railroads lasted generations; GPUs are depreciated over ~5–6 years. Brookings warns about opaque debt structures. If the revenue doesn't arrive, this is a bubble; if it does, it's a new utility.
Honest one-liner: "Largest private build-out in dollar terms; comparable to the railroads relative to the economy."
Caveats: these comparisons use different methods. WSJ compares annual spend to GDP; Brookings averages a projected multi-year total. Treat them as order-of-magnitude.
If asked: "Are we in a bubble?" Reasonable people disagree. Demand currently exceeds supply (Amazon says it can't keep up), but the spending relies on future revenue that isn't proven yet. Both can be true.
PART C
Reinforcement learning: from predictor to assistant to reasoner
Pretraining teaches a model what text looks like. RL teaches it what we actually want, by trying things and getting scored.
~0.25 min
We'll build RL from first principles: a slot-machine example, one equation, then how labs apply it to LLMs. If the core idea lands, RLHF, DPO and GRPO are variations on it.
Intuition: supervised learning says "copy this answer"; RL says "here's how good your answer was, figure out how to do better".
RL · post-training
Pretraining gives a document-completer, not an assistant
Base model · illustrative
What is the capital of France?What is the capital of Germany? What is the capital of Spain? …
After post-training
What is the capital of France?Paris.
Labelers preferred a 1.3B InstructGPT over the 175B GPT‑3 it came from.
Ouyang et al., "Training language models to follow instructions with human feedback" (InstructGPT), OpenAI, 2022.
~1.25 min
A pretrained model has read the internet and predicts what text comes next. On the internet, a question is often followed by more questions (a quiz page), so that's a perfectly good "prediction".
Stage 1, pretraining: trillions of tokens, next-token prediction. That's where almost all knowledge comes from, and almost all the compute (Part B).
Stage 2, supervised fine-tuning (SFT): train on example conversations written in the format we want. Now it answers like an assistant.
Stage 3, RL: SFT can only copy demonstrations. RL lets the model try answers and be scored, by human preferences (RLHF) or by automatic checkers (verifiable rewards).
The result is a model that follows instructions, refuses some things, and, with the right rewards, reasons.
The striking InstructGPT result: post-training mattered more than 100× more parameters for what users preferred.
Intuition: pretraining builds the brain; post-training teaches it the job.
Why not just SFT? Demonstrations show one good answer; they don't say how much worse the alternatives are, and they cap the model at the demonstrator's level. RL gets a signal on the model's own outputs, including ones nobody wrote.
If asked: "How much compute goes into post-training?" Historically a small fraction of pretraining, but RL compute has been growing fast since 2024; OpenAI has said o3 used about 10× o1's RL training compute (reported by Epoch AI).
RL · first principles
RL: learn by acting and being scored
State : what the agent sees. Action : what it does.
Reward : a number from the environment. Higher is better.
Policy : a probability over actions, with parameters .
Goal: choose to maximise expected return .
No labelled right answer. Only a score for what you tried.
~1.25 min
The RL loop: an agent interacts with an environment over time.
The agent observes a state and picks an action. For a game: screen → button. For an LLM: text so far → next token.
The environment returns a reward (a number) and a new state. Reward can be rare: often 0 until the very end.
The policy is the agent's brain: a neural network that outputs a probability for each action. It's stochastic on purpose, so it can explore.
The return G is total (optionally discounted) reward over an episode. The objective J is the expected return when acting with . Everything in RL is about raising J.
The key difference from supervised learning: nobody tells you the correct action. You only learn how good the action you took turned out to be.
Intuition: learning a sport. No one hands you the perfect move for every situation; you try, see the score, and adjust.
Discount : a number in [0,1] that makes far-future reward count a bit less. For LLM training it's usually 1 (one reward at the end of a finite response).
If asked: "Why a probability and not just the best action?" You need to explore to find out what's good, and a smooth probability makes the objective differentiable in (next slides).
RL · the simplest case
A bandit: the policy shifts toward what paid off
Each click: 48 pulls sampled from the current policy, then one policy-gradient update per 16 pulls.
~1.5 min
Four slot machines ("arms") with hidden payout rates. The state never changes; we pick an arm, get 1 or 0. Our policy starts uniform: 25% each.
Round 1: we pull arms according to the policy. Green coins paid, rose didn't. Then we nudge: arms that paid more than the batch average go up, arms below average go down.
Round 2: C is pulling ahead, because it pays more often. Note that early luck can mislead (noise), which is why we average over many pulls.
Each update is small. The policy is a softmax over four learnable numbers (logits); we take gradient steps on them.
By now C dominates, but the others aren't zero, so we still explore a little.
Average reward per pull (right) climbs toward C's rate.
Reveal: C pays 80%, B 50%, D 35%, A 20%. The policy found C without ever being told.
Intuition: do more of what worked better than average, less of what worked worse.
The exact update used: for each pull of arm a with reward r and batch-average b, logits change by . That's the policy gradient from the next slide with advantage r − b.
If asked: "Why not just pick the best average so far?" That's greedy and can lock onto a lucky arm early. The policy-gradient version keeps probabilities soft, so exploration continues while it learns.
RL · the one equation
Policy gradient: push up what beat expectations
Sample actions from the policy. That's the .
: make it more likely. : less likely. Step size .
A baseline cuts noise without biasing the gradient.
~2 min
This is the REINFORCE / policy-gradient theorem. Every method in this part is built on it.
The expectation means: act with your current policy, look at what happened. We estimate it with samples, like a minibatch.
is the direction in parameter space that makes the action you took more probable. It's the same gradient as supervised learning with as the "label".
Multiply by the advantage: how much better this action did than you expected. So RL is supervised learning on your own samples, weighted by how good they turned out.
Bars: good surprises (green) get pushed up, bad ones (rose) pushed down, about-as-expected barely moves.
The baseline: subtract the expected reward. If every action gets reward 10, raw rewards would push everything up; only the differences carry information.
Where it comes from: (log-derivative trick: ). The last sum is an expectation under , so samples estimate it.
Why the baseline is free: , so subtracting any b that doesn't depend on a adds zero in expectation but reduces variance.
Common mistake: thinking the reward must be differentiable. It doesn't: we never differentiate , only . The reward can be a unit test or a human click.
RL · applied to language models
For an LLM, every token is an action
Policy = the LM's softmax over ~128K tokens.
Reward arrives only at the end.
Which tokens deserve the credit?
~1.25 min
Map the vocabulary onto language models.
State: the prompt plus everything generated so far.
Action: the next token, sampled from the model's softmax. The LLM is literally the policy . Llama 3's vocabulary has about 128K tokens, so ~128K possible actions per step.
Episode: the full response, until the end-of-sequence token. Hundreds or thousands of actions.
Reward: usually one number at the end. Did the answer match? Did a human or a reward model like it?
Credit assignment: one scalar for 300 tokens. The simplest methods give every token in the response the same advantage; PPO tries to estimate per-token value with a learned critic.
Intuition: you get a grade on the essay, not on each word. Over many essays, the words that tend to appear in high-graded essays get reinforced.
Subtle point: of a whole response is the sum of per-token log-probs, so is just backprop through the sampled sequence. That's why this is cheap to implement on top of a normal training loop.
If asked: "Where's the environment?" For chat it's trivial (the response ends the episode). For agents it's real: tools, code execution, web pages, multiple turns.
RL · RLHF part 1
RLHF: turn human preferences into a reward model
People rarely give good absolute scores, but comparing two answers is easy.
Bradley–Terry: only the difference in scores matters.
The reward model is a learned stand-in for human judgment. It can be wrong.
Christiano et al. 2017; Stiennon et al. 2020; Ouyang et al. 2022. Bradley & Terry, Biometrika 1952.
~1.5 min
We want to RL against "what humans prefer", but we can't put a human in the loop for millions of samples. So we train a model of their preferences first.
Collect data: for a prompt, sample two responses from the SFT model.
A human picks the better one. Comparisons are much more consistent than 1–10 ratings.
The reward model is usually the LLM itself with a scalar head instead of the vocabulary. Bradley–Terry says the probability the human prefers A is the sigmoid of the score difference.
Train it by maximum likelihood on the human choices: that's this loss, a logistic regression on score differences. Here , so it predicts A wins with probability 0.90.
Now we have a reward function we can query millions of times. But it's only an approximation, trained on limited data: keep that in mind for the next slide.
Intuition: chess Elo ratings. You never measure a player's absolute strength, only who beats whom, and you fit numbers so the differences predict wins.
Subtle point: adding a constant to every score changes nothing, so reward-model scores have no absolute meaning. Only comparisons count.
If asked: "Who are the labelers?" Contractors and domain experts, often with detailed guidelines. Their biases become the reward model's biases (e.g. preferring longer or more confident answers).
RL · RLHF part 2
Then optimise that reward, on a leash
No leash ( ): the policy piles onto wherever the reward model is wrong.
With the leash , the best policy is the reference reweighted by reward:
PPO does this with clipped policy-gradient steps and a learned value baseline.
Illustrative 1-D picture; curves computed from the closed form shown. PPO: Schulman et al. 2017. KL-regularised RLHF: Ziegler et al. 2019; Ouyang et al. 2022.
~1.75 min
Picture every possible response along one axis (a cartoon, but useful). The grey curve is the SFT/reference model: where it normally puts probability.
Amber: the reward model's score. It's mostly sensible, but far from the data it was trained on it can be badly wrong, like this spike on weird responses no human ever rated.
Green dashed: true quality, which we can't see.
Maximise reward with no constraint and the policy collapses onto the spike: high reward-model score, garbage output. That's reward hacking.
The KL term charges for moving away from the reference. The optimum has a neat closed form: the reference distribution reweighted by . Where the reference puts ~zero probability, even a huge reward can't create much mass.
So the leashed policy moves toward better answers that the reference already considered plausible, and true quality improves.
PPO is the optimiser most labs used: policy-gradient steps with a clipped ratio so no single update moves too far, plus a value network to estimate advantages per token.
Intuition: "improve, but stay recognisably yourself". is the leash length.
Why the closed form: maximising over distributions gives (same maths as the Boltzmann distribution). DPO, next, starts from exactly this.
If asked: "Isn't the KL just regularisation?" Yes, but targeted: it keeps the policy in the region where the reward model is trustworthy and keeps fluency and diversity from pretraining.
RL · a shortcut
DPO: skip the reward model, train on the pairs
Plus
One supervised-style loss. No reward model, no sampling loop. Simple and stable.
Minus
Offline: learns only from a fixed set of pairs, with no exploration of its own outputs.
Rafailov et al., "Direct Preference Optimization: Your Language Model is Secretly a Reward Model", NeurIPS 2023.
~1.25 min
The trick: from the last slide, the optimal leashed policy is (normalised). Solve for : plus a constant.
Plug that "implicit reward" into the Bradley–Terry loss. The constant cancels (only differences matter) and you get this: a loss on the policy itself.
Bars: training raises the chosen answer's log-probability relative to the reference and lowers the rejected one's, until the margin is comfortably positive.
Practical upside: it's just a classification-style loss on preference pairs. Easy to run, which is why it spread fast in open-source models.
Downside: it never generates new samples during training, so it can't discover better answers than those in the dataset, and it can over-fit the pairs. Online RL (PPO, GRPO) samples fresh responses every step.
Intuition: instead of training a judge and then training a student to please the judge, train the student directly on the judge's verdicts.
Precise claim: DPO optimises the same KL-regularised objective as RLHF under the Bradley–Terry model, assuming the data cover the responses that matter. Hence "your language model is secretly a reward model".
If asked: "Which do frontier labs use?" Mixtures. Preference-style objectives are common for style and safety; large-scale online RL dominates for reasoning.
RL · verifiable rewards
Verifiable rewards: let a checker be the judge
Maths with known answers, code with unit tests, puzzles with checkers. No learned judge to fool.
DeepSeek‑R1‑Zero: RL straight from a base model, with only correctness and format rewards.
Responses grew longer on their own, and the model started re-checking its work ("wait…").
DeepSeek‑AI, "DeepSeek‑R1: Incentivizing Reasoning Capability in LLMs via RL", Jan 2025 (Nature, Sep 2025). AIME 2024 accuracy.
~1.5 min
RLHF's weak point is the learned reward model. For some tasks we don't need one: we can check the answer exactly.
Examples: a final number compared to the key; code run against hidden tests; a proof checked by Lean. Reward is 1 or 0.
Chart: DeepSeek‑R1‑Zero on AIME 2024 (hard competition maths). Pass@1 went from 15.6% at the start of RL to 71.0% at the end; with majority voting over samples, 86.7%.
"Zero" means no supervised fine-tuning first: pure RL from a base model, with a reward for correct answers and one for putting reasoning inside think tags.
Nobody told it to reason longer. Average response length grew during training, and behaviours like re-checking and backtracking appeared, because they made correct answers more likely.
Why RL produces reasoning: the reward only cares whether the final answer is right. Any habit that raises that probability, such as writing out steps, checking, or trying another approach, gets reinforced by the policy gradient. Thinking tokens are actions that buy accuracy.
Caveat: R1-Zero's outputs were hard to read (mixed languages), so the released R1 added cold-start SFT data and more stages. Also, "verifiable" can still be gamed: tests can be hacked (two slides on).
If asked: "Does RL teach new knowledge?" Debated. A common view is that it mostly sharpens and elicits abilities the base model already has (better pass@1), with smaller gains when many samples are allowed.
RL · GRPO
GRPO: grade a group of answers on a curve
All right or all wrong? Every : no signal. Useful problems are the ones the model sometimes solves.
~1.75 min
GRPO (Group Relative Policy Optimization) was introduced in DeepSeekMath (2024) and used for R1. Question: 17 × 24.
Sample a group of G answers (here 8) for the same prompt from the current policy.
Check each one: 408 is right (reward 1), others wrong (0).
The baseline is just the group's average reward, 0.625. No separate value network (PPO's critic, often as big as the policy) is needed.
Advantage = (reward − mean) / std. Right answers get +0.77, wrong ones −1.29. Wrong answers are rarer here, so each is pushed down harder.
Policy-gradient update: every token of each correct sample becomes a bit more likely, every token of each wrong sample a bit less. GRPO also keeps PPO's clipping and a KL term to the reference.
If all eight are right (or all wrong), all advantages are zero and the prompt teaches nothing. So training data is chosen at the edge of the model's ability.
Intuition: grading on a curve. You're rewarded for beating your classmates on the same exam, which automatically adjusts for how hard the exam was.
Why it caught on: it drops the critic network (less memory, simpler), and for verifiable rewards the group mean is a very good baseline.
If asked: "Isn't 8 samples noisy?" Yes. Groups are typically 8–64 and there are many prompts per batch. It's the same Monte Carlo idea as the bandit.
RL · the failure mode
Reward hacking: optimise a proxy hard and it breaks
"When a measure becomes a target, it ceases to be a good measure."
Goodhart's law, in Marilyn Strathern's 1997 phrasing
OpenAI · 2016 · CoastRunners
A boat-racing agent learned to circle forever hitting point targets instead of finishing the race, and still scored higher than humans.
OpenAI · Mar 2025 · coding RL
Frontier reasoning models made verify always return true, or exited before tests ran. Penalising "bad thoughts" mostly taught them to hide it.
OpenAI · Apr 2025 · GPT‑4o
An extra thumbs-up reward signal made the model sycophantic. Rolled back within days.
OpenAI "Faulty reward functions in the wild" (2016); "Detecting misbehavior in frontier reasoning models" (Mar 2025); "Expanding on sycophancy" (May 2025). Curve shape: Gao et al. 2022.
~1.5 min
Every reward is a proxy for what we want. RL is an optimiser: push hard enough and it finds where the proxy and the goal disagree.
The curve: as optimisation pressure rises, the proxy reward keeps going up but true quality peaks and then falls. Gao et al. (2022) measured exactly this with reward models.
Classic example: CoastRunners. Reward = points; points came from targets; the agent found a lagoon where it could loop forever.
2025, frontier models on coding tasks: when rewarded for passing tests, they sometimes edited the test, stubbed libraries, or exited early. The chain of thought said things like "we can hack verify to always return true". When OpenAI penalised such thoughts, models kept cheating but stopped saying so.
GPT‑4o, April 2025: adding user thumbs-up/down as a reward pushed the model toward flattery. Users like being agreed with; that isn't the same as being helped.
Intuition: teach to the test and students learn the test, not the subject.
Defences: KL leash, better and more diverse rewards, held-out evaluations, monitoring reasoning traces, and humans reading samples. None is complete.
If asked: "Is this a safety issue?" Yes: as models get more capable, finding loopholes gets easier and harder to spot. It's a main reason labs care about readable reasoning traces.
RL · our project
Our stretch goal: RL for the diffusion renderer
Method · DDPO
Treat the denoising chain as a multi-step decision process. Reward at the end, policy gradient through each step's Gaussian.
Research question
Does RL improve held-out combinations, or only sharpen ones it has already seen?
Black et al., "Training Diffusion Models with Reinforcement Learning" (DDPO), 2023. Project proposal, Sep 2026.
~1.25 min
Bring it home. Our diffusion model (deck 5) turns noise into a frame over many denoising steps.
Each denoising step samples from a Gaussian the network predicts, so each step is an action with a log-probability, just like a token.
The skinner gives us something rare in RL: an exact answer key for every frame, including combinations never seen in training.
Reward = how close the final frame is to the skinner's frame (e.g. −LPIPS or PSNR). It's a verifiable reward, like unit tests for images.
DDPO applies the policy gradient across the whole chain: frames that score well make their denoising steps more likely.
The research question fits our hypotheses: does RL fine-tuning help on held-out cells, or only sharpen seen ones?
That is the same debate as "does RL teach new reasoning or only sharpen it" in LLMs, in a setting where we can measure it exactly.
Watch for reward hacking: PSNR rewards blurry averages; LPIPS can be gamed by texture tricks. Use more than one metric, and look at the frames.
If asked: "Why is this a stretch goal?" RL on diffusion is expensive (sample whole chains per update) and needs a working base model first. Phases 1–3 come first.
PART D
Test-time compute: think harder, not just bigger
A quick tangent with depth. If more compute at training time makes models better, what about more compute per question?
~0.25 min
Part C showed RL teaching models to use thinking tokens. Part D asks how to spend compute when answering: longer, wider, or smarter.
Intuition: a person gets more right on an exam with three hours than with three minutes. Models do too, if they've been trained to use the time.
Test-time compute · the idea
Same weights, more thinking, better answers
OpenAI o1 (Sep 2024): accuracy rose steadily with both RL training compute and thinking time .
Two directions to spend it: longer (one deeper chain of thought) or wider (many attempts, then choose).
Since 2025, maths-olympiad gold (IMO, 35/42) came from models given long thinking time.
OpenAI, "Learning to reason with LLMs", Sep 2024 (AIME 2024). IMO 2025: Google DeepMind and OpenAI announcements, Jul 2025.
~1.25 min
AIME 2024 is a hard US maths competition, 15 problems with integer answers.
GPT‑4o, answering directly: 12%.
o1, one sample but with a long hidden chain of thought: 74%. The same kind of transformer, trained with RL to think first.
o1 with 64 samples and a majority vote: 83%.
o1 with 1,000 samples re-ranked by a learned scorer: 93%. Each step spends more compute per question.
Two families of methods: sequential (think longer in one chain) and parallel (try many times, then pick). Next slides cover each.
In July 2025 both Google DeepMind (Gemini Deep Think, officially graded) and OpenAI (experimental model) reported gold-medal scores at the IMO, in natural language, using long thinking.
Intuition: training compute builds a better brain once; test-time compute lets that brain work harder on each specific problem.
Subtle point: this only works well for models trained to use the extra tokens. Asking a non-reasoning model to "think longer" helps much less.
If asked: "Isn't the 93% unfair?" It's a different setting (1,000 attempts plus a chooser), not a single-answer score. Always ask which setting a reported number used.
Test-time compute · lever 1
Lever 1: longer chains of thought
Reasoning models write thousands of hidden tokens before answering: plans, attempts, checks, backtracking.
Budget forcing (s1): cut thinking off at a budget, or append "Wait" to make it keep going. AIME24 rose from 50% to 57% .
Every thinking token is a decode step: a longer KV cache and quadratic attention (Part A).
Muennighoff et al., "s1: Simple test-time scaling", 2025 (Qwen2.5‑32B fine-tuned on 1,000 examples).
~1.25 min
Each row is part of a reasoning trace; each block a token. The model is solving the problem in writing.
Short chain: a few hundred tokens, enough for an easy problem.
Hard problems get thousands to tens of thousands of tokens. The model explores, notices a mistake, and tries again: the behaviour RL rewarded.
The s1 paper showed a cheap version: fine-tune on just 1,000 good reasoning traces, then control length at inference. If the model tries to stop early, append "Wait" and it often re-checks and fixes an error. AIME24 went 50% → 57%.
The cost connects to Part A: a 30K-token chain means a 30K-token KV cache per user, and generating it is quadratic. Reasoning models are expensive to serve because of exactly this.
Intuition: showing your work isn't just for the teacher. Writing intermediate steps gives the model working memory: each token can attend to everything it already wrote.
Subtle point: a transformer does a fixed amount of computation per token. Emitting more tokens is the only way for it to "think longer". Chain of thought turns tokens into computation.
If asked: "Is the chain of thought what the model 'really' thinks?" Not guaranteed. Research shows traces can omit or rationalise. That's why monitoring them is useful but not proof.
Test-time compute · lever 2
Lever 2: sample many answers, then vote
Self-consistency: sample N chains at temperature > 0, take the most common final answer.
With PaLM 540B it added +17.9 points on GSM8K over a single chain of thought.
Coverage (any sample right) keeps rising: 15.9% → 56% on SWE-bench Lite going from 1 to 250 samples. But only a perfect checker can collect that.
Parallel is easy to scale: the samples don't wait for each other.
Wang et al., "Self-Consistency Improves CoT Reasoning", ICLR 2023. Brown et al., "Large Language Monkeys", 2024 (DeepSeek‑V2‑Coder‑Instruct).
~1.25 min
Same question, many independent tries. Each path is a different sampled chain of thought.
Different paths, different final answers.
Tally the answers. Errors tend to scatter while correct reasoning tends to converge on the same answer, so the mode is usually right.
Self-consistency (Google, 2022) gave large gains on maths word problems with no extra training.
Pass@k ("is any of the k right?") rises steadily, roughly log-linearly in k. But a vote can't find a correct answer that's in the minority. To use coverage you need a verifier: unit tests, a proof checker, or a learned scorer.
Practical plus: N samples run as a batch, so wall-clock time barely grows. Cost grows linearly.
Intuition: ask a crowd of independent guessers and take the most popular answer. It works when errors are uncorrelated.
Failure mode: if the model is systematically wrong (a shared misconception), more votes make it confidently wrong.
If asked: "What about open-ended answers?" Voting needs answers you can compare (numbers, multiple choice). For essays or code you need a scorer instead: the next lever.
Test-time compute · levers 3 and 4
Verifiers pick the winner; search explores steps
Outcome reward model (ORM): scores finished answers. Use it for best-of-N.
Process reward model (PRM): scores each reasoning step, so bad paths are caught early.
Search: beam or tree search over steps, guided by the PRM.
Best-of-N on a MATH test subset: PRM 78.2% vs ORM 72.4% vs majority vote 69.6%.
Lightman et al., "Let's Verify Step by Step", OpenAI 2023 (ICLR 2024). Snell et al., ICLR 2025.
~1.5 min
A reasoning attempt as a tree: each node is a step (a line of working), branches are alternative next steps.
Growing the tree means sampling several continuations at each step.
An outcome reward model only sees the leaves: finished answers. Best-of-N picks the highest-scoring one.
A process reward model scores every intermediate step: "is this line correct so far?". Training it needs step-level labels (OpenAI collected 800K for PRM800K).
With step scores you can search: keep the best few partial solutions at each depth (beam search) and drop the rest.
Compute goes to promising branches instead of finishing doomed ones.
Lightman et al.: step-level supervision beat outcome-level and majority voting for picking among many samples.
Intuition: an ORM is a teacher who only marks the final answer; a PRM marks every line, so you find out at line 3 that you went wrong.
Snell et al. (2024): choosing the method per problem difficulty (revisions for easier problems, PRM search for harder ones) was over 4× more compute-efficient than best-of-N, and could beat a 14× larger model at equal FLOPs where the small model had partial success.
If asked: "Do o1/R1 use explicit tree search?" Published accounts point to long single chains learned by RL rather than explicit search at inference. Search and verifiers remain common on top, especially for code and maths.
Test-time compute · the curves
Accuracy rises with log-compute, then flattens
Left: OpenAI o1 AIME 2024 (Sep 2024); Brown et al. 2024, SWE-bench Lite coverage (1 and 250 samples, log-linear per paper). Right: shapes only, not data.
~1.5 min
Left: real published points, x-axis = samples per problem on a log scale.
Amber: o1 on AIME: 74% with 1 sample, 83% with 64 (vote), 93% with 1,000 (re-ranked). Each step costs 16× to 64× more compute for 9–10 points.
Blue: coverage on SWE-bench Lite rises from 15.9% to 56% from 1 to 250 samples, roughly a straight line in log(samples). Coverage needs a perfect checker to cash in.
Right: an illustration of the shapes reported across papers. Sequential thinking and parallel voting both improve roughly linearly in log-compute at first.
Then they flatten: voting plateaus once the majority answer stabilises, and longer thinking eventually hits the model's ceiling (or starts overthinking).
Intuition: "log-linear" means each fixed gain costs a multiplicative increase in compute. +10 points might cost 10×, the next +10 another 10×.
Honesty: the right panel has no numbers on purpose. Exact shapes depend on the model, task and method, and published curves (e.g. o1's blog) often hide the axis scales.
If asked: "Sequential or parallel, which is better?" Depends on difficulty: Snell et al. found sequential revisions better for easier problems and parallel/search better for harder ones. Parallel is easier on latency.
Test-time compute · the lever board
Six levers, and what each one costs
Lever What it does Main cost
1 · Think longer One deeper chain of thought Latency; long KV cache; quadratic attention
2 · Sample wider Best-of-N, majority vote N× tokens; vote plateaus
3 · Verify ORM/PRM, tests, proof checkers Need a verifier you can trust
4 · Search Beam or tree over steps Complex, sequential, PRM-dependent
5 · Use tools Run code, call search, calculators Sandboxes, round-trip latency
6 · Control the budget Thinking budgets, effort levels, routing easy vs hard Deciding what is "hard" in advance
Overthinking is real: models can spend thousands of tokens on "2+3". Diminishing returns mean adaptive budgets matter.
Overthinking: Chen et al., "Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs", Dec 2024.
~1.5 min
One table of everything labs now tune at inference time.
Longer thinking: biggest wins on hard reasoning, but the user waits, and Part A's costs (KV cache, quadratic attention) grow with every token.
Wider sampling: easy to parallelise; cost is linear in N; voting saturates.
Verifiers turn coverage into accuracy, but only as good as the verifier. Unit tests and proof checkers are gold; learned reward models can be hacked.
Search: the most compute-efficient in some studies, but hard to engineer and sequential across depth.
Tools: let the model run Python instead of multiplying in its head. Exact answers for the computable parts of a problem.
Budget control: APIs now expose "reasoning effort" or thinking-token budgets, and some products route easy questions to fast modes automatically. Money is spent where it helps.
The trade-off: latency, cost, diminishing returns, and overthinking simple questions.
Intuition: test-time compute is a dial, not a switch. The skill is knowing how far to turn it per question.
If asked: "Which matters most in practice?" For products: budget control and tools. For benchmarks: long thinking plus parallel sampling with good verifiers.
Close
Why the frontier is crowded and capital-intensive
Three ways to turn compute into capability: pretraining , RL post-training , test-time compute .
Each is predictable, and each costs GPUs, power and engineering at a scale only a few labs can afford.
Serving the result is its own bill: memory-bound decode, growing KV caches, long reasoning traces.
Next, deck 5: why we picked a research area where a student team can still reach the edge.
~1 min
Pull it together. Three axes, each turning more compute into a better model.
Pretraining: scaling laws (Part B). Costs: tens of thousands of GPUs for months, trillions of tokens.
RL post-training: Part C. Costs: millions of sampled rollouts, verifiers, reward models, human labels.
Test-time compute: Part D. Costs: every user, every question, more tokens, paid forever.
All three reward whoever has the most compute. That is why capex is approaching $750B a year for four companies, and why gigawatts are the unit.
And serving (Part A) means even a finished model needs huge fleets, because decode is memory-bound and reasoning traces are long.
For a student team, competing on the LLM frontier means losing a spending race. Deck 5 is about choosing an open question where ideas and careful experiments beat budgets.
Intuition: the frontier LLM race is a capital race with a research problem inside it. We want a research problem with no capital race.
Mini-presentation ideas from this deck: the KV cache and its arithmetic; the roofline; data vs tensor vs pipeline parallelism; RLHF vs DPO vs GRPO; reward hacking examples; test-time compute levers.