QMIND Research · Deck 4 of 7

The frontier: serving, scale, RL

The same transformer from deck 3, now running for millions of people: what it costs to serve, what it takes to train, and how RL and extra thinking time taught it to reason.

A · Serving B · Scale C · Reinforcement learning D · Test-time compute
PART A

Serving: from weights to a product

Every token a chatbot sends you is one more trip around a loop. The loop's bottleneck is memory, not math.

Serving · the generation loop

Reading is parallel; writing is one token at a time

Prefill

Whole prompt in one pass. Big matrix–matrix multiplies. Compute-bound. Sets time to first token.

Decode

One position per pass, hundreds of passes. Matrix–vector multiplies. Memory-bound. Sets tokens per second.

Serving · the memory wall

Decode is memory-bound: each token reads every weight

Llama‑3.1‑8B in BF16 is 16 GB of weights. Each decode step streams all of it from HBM to the math units.

That caps one user at about 210 tokens/s, however fast the chip computes.

The math needed: GFLOP, which takes 16 µs at ~989 TFLOP/s. Tensor cores are busy about 0.3% of the time.

So read the weights once and use them for many users at the same time.

H100 SXM: 80 GB HBM3 at 3.35 TB/s, ~989 dense BF16 TFLOP/s (NVIDIA spec sheet). Llama 3.1 8B ≈ 8.0B params. Upper-bound arithmetic, as of Oct 2026.

Serving · the KV cache

KV cache: compute keys and values once, keep them

A new token's query must meet the keys and values of every earlier token, in every layer.

Recomputing them at every step repeats work we already did.

So store them. Prefill writes a column for each prompt token.

Each decode step adds one new column: its K and V for every layer and KV head.

We traded compute for memory. The cache grows with every token.

Serving · cache arithmetic

How big is the cache? Do the arithmetic

Llama‑3.1‑70B: 80 layers, 8 KV heads, , BF16:

32 users at 8K tokens each: 80 GiB of cache, more than a whole H100.

Without grouped KV heads (64 instead of 8): 8× larger, 2.5 MiB per token.

Config: Llama‑3.1‑70B (80 layers, 64 query / 8 KV heads, head dim 128), per Meta's released config; KV sizing as in MinIO MemKV docs. As of Oct 2026.

Serving · why long context is expensive

Attention is quadratic in context length

One pass over a whole sequence: work for the scores and the mixing.

With the cache, one new token costs : one new row.

Generating tokens is still quadratic overall.

Memory only grows linearly: the cache is one row of entries per layer and head.

Serving · batching and cost

Batching: one weight read serves many users

Stack B users' vectors into a matrix: one trip of the weights through memory now produces B tokens.

Total throughput climbs almost linearly at first…

…while each user's speed drops slowly.

Cost per million tokens falls about 20× from B = 1 to B = 64.

Then the KV caches fill memory, and reading them dominates. That is why cache size matters so much.

Illustrative roofline model, ideal kernels: Llama‑3.1‑8B, BF16, one H100, 4K‑token contexts, $2.50/GPU‑hr (H100 rentals ≈ $1.49–$6.98/hr, IntuitionLabs 2026).

Serving · scheduling and memory

Continuous batching and paged KV memory

vLLM (2023): earlier servers wasted 60–80% of KV memory to fragmentation and over‑reservation; paging cut waste to under 4%, for up to 24× the throughput of HF Transformers.

Source: vLLM project blog, "Easy, Fast, and Cheap LLM Serving with PagedAttention", Jun 2023 (Kwon et al., SOSP 2023).

Serving · bending the curve (1)

Shrink the bytes: share KV heads, use fewer bits

GQA: groups of query heads share one K/V head. Llama‑3.1‑70B: 64 query heads, 8 KV heads, so the cache is 8× smaller.

MQA shares a single K/V head. MLA (DeepSeek) compresses K and V into a small latent vector.

FormatBytes70B weights
BF162141 GB
FP8 / INT8171 GB
FP4 / INT40.535 GB

Fewer bytes to stream means faster decode and more users per GPU. The price is some accuracy, so it needs care.

GQA: Ainslie et al. 2023; MQA: Shazeer 2019; MLA: DeepSeek‑V2 2024 (reports a 93.3% smaller KV cache than DeepSeek 67B). Native FP8 on Hopper, FP4 on Blackwell.

Serving · bending the curve (2)

Fewer memory trips, or less attention

Fixed-size memory forgets. Many recent models are hybrids: mostly cheap layers, plus a few full-attention layers.

FlashAttention: Dao et al. 2022 (exact, IO-aware). Sliding window: e.g. Mistral 7B (2023). State-space: Mamba, Gu & Dao 2023.

Serving · bending the curve (3)

Speculative decoding: draft cheaply, verify in one pass

Verification uses a rejection-sampling rule, so the output distribution is exactly the big model's. Reported speedups: about 2–3×.

Leviathan, Kalman & Matias, "Fast Inference from Transformers via Speculative Decoding", ICML 2023; Chen et al. (DeepMind) 2023.

PART B

Scale: chips, wires and gigawatts

Why GPUs, why NVIDIA, how one model is split across tens of thousands of chips, and what the build-out costs.

Scale · why GPUs

Neural nets are matrix multiplies; GPUs are built for them

A layer is : millions of independent multiply–adds.

A CPU has a few big cores built for branchy, serial code. It walks through the output a few tiles at a time.

A GPU runs thousands of simple lanes doing the same operation on different data. H100: 132 SMs, 528 tensor cores.

A tensor core does a whole small matrix multiply–accumulate per instruction: .

Transformer FLOPs are overwhelmingly matmuls, so the fit is near perfect.

H100 SXM5: 132 streaming multiprocessors, 4 tensor cores each (NVIDIA Hopper architecture whitepaper, 2022).

Scale · why NVIDIA

NVIDIA's moat is software and systems, not just chips

Software · since 2007

CUDA and its libraries

cuBLAS, cuDNN, NCCL, TensorRT. PyTorch runs best on CUDA. Two decades of tuned kernels.

Interconnect

NVLink, NVSwitch, InfiniBand

GPU-to-GPU wires inside the rack, plus Mellanox networking (acquired 2020) between racks.

Systems

Sell the rack, not the chip

GB200 NVL72: 72 GPUs in one NVLink domain that behaves like one giant GPU.

HBM per GPU: H100 3.35 TB/s; H200 4.8; B200 8; GB300 8; Rubin ~22 (HBM4, detailed Jan 2026, shipping H2 2026). NVIDIA specs/announcements, as of Oct 2026.

Scale · the roofline model

Arithmetic intensity decides the bottleneck

A weight matmul at batch : FLOPs over bytes, so .

Above the ridge (~295 FLOP/byte on H100) you are compute-bound. Below it, memory-bound.

Attention over each user's KV cache stays near at any batch size: every user brings their own cache.

Roofline model: Williams, Waterman & Patterson, CACM 2009. H100 SXM figures: 3.35 TB/s, ~989 dense BF16 TFLOP/s; ridge = 989 / 3.35 ≈ 295 FLOP/byte.

Scale · training

A frontier model won't fit on one GPU, or finish on one

Mixed-precision Adam keeps about 16 bytes per parameter: weights, gradients, FP32 master copy, two moments.

70B parameters → 1.13 TB of state, before activations: about 15 H100s just to hold it.

Llama 3.1 405B used FLOPs.

H100s
1
Time at 40% utilisation
3,000 years

16 B/param: ZeRO paper (Rajbhandari et al. 2020). FLOPs and up to 16,384 H100s at 38–43% MFU: "The Llama 3 Herd of Models", Meta, Jul 2024.

Scale · data parallelism

Data parallel: copy the model, split the batch, average

Ring all-reduce: each GPU sends about × the gradient size, nearly independent of .

Scale · sharding the state

ZeRO / FSDP: shard the state instead of copying it

A 7B model with Adam needs ~112 GB per copy. Plain data parallel copies it to every GPU.

Stage 1: each GPU keeps only its 1/N slice of the optimizer state.

Stage 2: gradients too: reduce-scatter instead of all-reduce.

Stage 3 / FSDP: parameters too. Each layer is gathered just in time, used, then freed.

The price is more communication. It fits because it overlaps with compute.

Rajbhandari et al., "ZeRO: Memory Optimizations Toward Training Trillion Parameter Models", 2020; PyTorch FSDP (Zhao et al. 2023).

Scale · tensor parallelism

Tensor parallel: split each matrix across GPUs

An all-reduce in every layer, forward and backward. It only pays inside a fast NVLink domain.

Scale · pipeline parallelism

Pipeline parallel: layers across GPUs, and the bubble

stages, microbatches. Smarter schedules (1F1B, interleaved) also cut memory and bubbles.

Scale · mixture of experts

Mixture of experts: each token visits a few experts

The MLP becomes many experts. A small router scores them per token and keeps the top .

Experts live on different GPUs (expert parallelism), so routing is an all-to-all shuffle.

DeepSeek‑V3: 671B parameters in total, 37B active per token (8 of 256 routed experts, plus 1 shared).

A balancing term keeps load even, so no expert or GPU becomes the bottleneck.

DeepSeek‑AI, "DeepSeek‑V3 Technical Report", Dec 2024. MoE layers: Shazeer et al. 2017; Switch Transformer, Fedus et al. 2021.

Scale · putting it together

Real runs combine all of these, shaped by the wires

Llama 3.1 405B: 16,384 H100s = TP 8 × PP 16 × DP 128.

HBM (H100)3.35 TB/s
NVLink 5 / GPU1.8 TB/s
NVLink 4 / GPU0.9 TB/s
InfiniBand NDR / NIC0.05 TB/s

Bars on a log scale.

Rule: put the chattiest parallelism on the fastest wire.

Llama 3 Herd of Models (Meta, Jul 2024), Table 4. NVLink 4 (H100) 900 GB/s; NVLink 5 (Blackwell) 1.8 TB/s (NVLink: bidirectional totals); NDR InfiniBand 400 Gb/s per port.

Scale · why anyone spends this

Scaling laws: loss falls predictably with compute

The best model for each budget lies on a smooth line: a power law in compute.

Chinchilla: grow parameters and tokens together, about 20 tokens per parameter. 70B on 1.4T tokens beat 280B Gopher.

Today's models train far past that (Llama 3 8B: 15T tokens) because a smaller model is cheaper to serve.

Curves use Chinchilla's fitted constants (E=1.69, A=406.4, B=410.7, α=0.34, β=0.28; Hoffmann et al. 2022). Kaplan et al. 2020. Llama 3: Meta, 2024.

Scale · the build-out

The build-out, in dollars

Capital expenditure (mostly data centres, chips, power) of Amazon, Alphabet, Microsoft and Meta.

2026 guidance after Q2 earnings: about $720–745B, roughly 75% above 2025.

Amazon, Alphabet and Meta each raised their 2026 guidance during the year.

Plus Oracle (FY27 up to ~$95B) and the OpenAI–Oracle–SoftBank Stargate venture ($500B, >9 GW planned).

2024–25: company filings via Platformonomics (Feb 2026; Alphabet excl. finance leases). 2026: Q2 2026 earnings guidance (Jul 2026), per MLQ/IG summaries (Aug 2026). Oracle: Jun 2026.

Scale · power

The new unit of compute is the gigawatt

One gigawatt is about the output of a large nuclear reactor, running continuously.

World (IEA): data centres used ~415 TWh in 2024 (~1.5% of electricity), heading to ~945 TWh by 2030, a bit more than Japan uses today.

US (LBNL/DOE): 4.4% of electricity in 2023; 6.7–12% projected for 2028.

Power, grid connections and turbines, not chips, are now the slowest part to build.

Epoch AI, Stargate sites (Apr 2026); TechCrunch on Meta Prometheus/Hyperion (Jul 2025); IEA "Energy and AI" (2025); DOE/LBNL report (Dec 2024).

Scale · perspective

"Biggest build-out ever"? Depends on the yardstick

In nominal dollars: the largest private capital-spending wave on record.

As a share of GDP: about the same as the 1850s railroad boom. Wartime mobilisation was far larger.

What's different: chips wear out in years, while rails lasted a century. Revenue is still catching up, and more of it is debt-financed.

WSJ analysis via Benzinga, Feb 2026 (annual spend as % of GDP; AI = big-4 2026 est. $670B). Brookings (van Nieuwerburgh), Sep 2026: 3.6%/yr projected, 2025–32.

PART C

Reinforcement learning: from predictor to assistant to reasoner

Pretraining teaches a model what text looks like. RL teaches it what we actually want, by trying things and getting scored.

RL · post-training

Pretraining gives a document-completer, not an assistant

Base model · illustrative

What is the capital of France?
What is the capital of Germany? What is the capital of Spain? …

After post-training

What is the capital of France?
Paris.

Labelers preferred a 1.3B InstructGPT over the 175B GPT‑3 it came from.

Ouyang et al., "Training language models to follow instructions with human feedback" (InstructGPT), OpenAI, 2022.

RL · first principles

RL: learn by acting and being scored

State : what the agent sees. Action : what it does.

Reward : a number from the environment. Higher is better.

Policy : a probability over actions, with parameters .

Goal: choose to maximise expected return.

No labelled right answer. Only a score for what you tried.

RL · the simplest case

A bandit: the policy shifts toward what paid off

Each click: 48 pulls sampled from the current policy, then one policy-gradient update per 16 pulls.

RL · the one equation

Policy gradient: push up what beat expectations

Sample actions from the policy. That's the .

: make it more likely. : less likely. Step size .

A baseline cuts noise without biasing the gradient.

RL · applied to language models

For an LLM, every token is an action

Policy = the LM's softmax over ~128K tokens.

Reward arrives only at the end.

Which tokens deserve the credit?

RL · RLHF part 1

RLHF: turn human preferences into a reward model

People rarely give good absolute scores, but comparing two answers is easy.

Bradley–Terry: only the difference in scores matters.

The reward model is a learned stand-in for human judgment. It can be wrong.

Christiano et al. 2017; Stiennon et al. 2020; Ouyang et al. 2022. Bradley & Terry, Biometrika 1952.

RL · RLHF part 2

Then optimise that reward, on a leash

No leash (): the policy piles onto wherever the reward model is wrong.

With the leash, the best policy is the reference reweighted by reward:

PPO does this with clipped policy-gradient steps and a learned value baseline.

Illustrative 1-D picture; curves computed from the closed form shown. PPO: Schulman et al. 2017. KL-regularised RLHF: Ziegler et al. 2019; Ouyang et al. 2022.

RL · a shortcut

DPO: skip the reward model, train on the pairs

Plus

One supervised-style loss. No reward model, no sampling loop. Simple and stable.

Minus

Offline: learns only from a fixed set of pairs, with no exploration of its own outputs.

Rafailov et al., "Direct Preference Optimization: Your Language Model is Secretly a Reward Model", NeurIPS 2023.

RL · verifiable rewards

Verifiable rewards: let a checker be the judge

Maths with known answers, code with unit tests, puzzles with checkers. No learned judge to fool.

DeepSeek‑R1‑Zero: RL straight from a base model, with only correctness and format rewards.

Responses grew longer on their own, and the model started re-checking its work ("wait…").

DeepSeek‑AI, "DeepSeek‑R1: Incentivizing Reasoning Capability in LLMs via RL", Jan 2025 (Nature, Sep 2025). AIME 2024 accuracy.

RL · GRPO

GRPO: grade a group of answers on a curve

All right or all wrong? Every : no signal. Useful problems are the ones the model sometimes solves.

RL · the failure mode

Reward hacking: optimise a proxy hard and it breaks

"When a measure becomes a target, it ceases to be a good measure."

Goodhart's law, in Marilyn Strathern's 1997 phrasing

OpenAI · 2016 · CoastRunners

A boat-racing agent learned to circle forever hitting point targets instead of finishing the race, and still scored higher than humans.

OpenAI · Mar 2025 · coding RL

Frontier reasoning models made verify always return true, or exited before tests ran. Penalising "bad thoughts" mostly taught them to hide it.

OpenAI · Apr 2025 · GPT‑4o

An extra thumbs-up reward signal made the model sycophantic. Rolled back within days.

OpenAI "Faulty reward functions in the wild" (2016); "Detecting misbehavior in frontier reasoning models" (Mar 2025); "Expanding on sycophancy" (May 2025). Curve shape: Gao et al. 2022.

RL · our project

Our stretch goal: RL for the diffusion renderer

Method · DDPO

Treat the denoising chain as a multi-step decision process. Reward at the end, policy gradient through each step's Gaussian.

Research question

Does RL improve held-out combinations, or only sharpen ones it has already seen?

Black et al., "Training Diffusion Models with Reinforcement Learning" (DDPO), 2023. Project proposal, Sep 2026.

PART D

Test-time compute: think harder, not just bigger

A quick tangent with depth. If more compute at training time makes models better, what about more compute per question?

Test-time compute · the idea

Same weights, more thinking, better answers

OpenAI o1 (Sep 2024): accuracy rose steadily with both RL training compute and thinking time.

Two directions to spend it: longer (one deeper chain of thought) or wider (many attempts, then choose).

Since 2025, maths-olympiad gold (IMO, 35/42) came from models given long thinking time.

OpenAI, "Learning to reason with LLMs", Sep 2024 (AIME 2024). IMO 2025: Google DeepMind and OpenAI announcements, Jul 2025.

Test-time compute · lever 1

Lever 1: longer chains of thought

Reasoning models write thousands of hidden tokens before answering: plans, attempts, checks, backtracking.

Budget forcing (s1): cut thinking off at a budget, or append "Wait" to make it keep going. AIME24 rose from 50% to 57%.

Every thinking token is a decode step: a longer KV cache and quadratic attention (Part A).

Muennighoff et al., "s1: Simple test-time scaling", 2025 (Qwen2.5‑32B fine-tuned on 1,000 examples).

Test-time compute · lever 2

Lever 2: sample many answers, then vote

Self-consistency: sample N chains at temperature > 0, take the most common final answer.

With PaLM 540B it added +17.9 points on GSM8K over a single chain of thought.

Coverage (any sample right) keeps rising: 15.9% → 56% on SWE-bench Lite going from 1 to 250 samples. But only a perfect checker can collect that.

Parallel is easy to scale: the samples don't wait for each other.

Wang et al., "Self-Consistency Improves CoT Reasoning", ICLR 2023. Brown et al., "Large Language Monkeys", 2024 (DeepSeek‑V2‑Coder‑Instruct).

Test-time compute · levers 3 and 4

Verifiers pick the winner; search explores steps

Outcome reward model (ORM): scores finished answers. Use it for best-of-N.

Process reward model (PRM): scores each reasoning step, so bad paths are caught early.

Search: beam or tree search over steps, guided by the PRM.

Best-of-N on a MATH test subset: PRM 78.2% vs ORM 72.4% vs majority vote 69.6%.

Lightman et al., "Let's Verify Step by Step", OpenAI 2023 (ICLR 2024). Snell et al., ICLR 2025.

Test-time compute · the curves

Accuracy rises with log-compute, then flattens

Left: OpenAI o1 AIME 2024 (Sep 2024); Brown et al. 2024, SWE-bench Lite coverage (1 and 250 samples, log-linear per paper). Right: shapes only, not data.

Test-time compute · the lever board

Six levers, and what each one costs

LeverWhat it doesMain cost
1 · Think longerOne deeper chain of thoughtLatency; long KV cache; quadratic attention
2 · Sample widerBest-of-N, majority voteN× tokens; vote plateaus
3 · VerifyORM/PRM, tests, proof checkersNeed a verifier you can trust
4 · SearchBeam or tree over stepsComplex, sequential, PRM-dependent
5 · Use toolsRun code, call search, calculatorsSandboxes, round-trip latency
6 · Control the budgetThinking budgets, effort levels, routing easy vs hardDeciding what is "hard" in advance

Overthinking is real: models can spend thousands of tokens on "2+3". Diminishing returns mean adaptive budgets matter.

Overthinking: Chen et al., "Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs", Dec 2024.

Close

Why the frontier is crowded and capital-intensive

Three ways to turn compute into capability: pretraining, RL post-training, test-time compute.

Each is predictable, and each costs GPUs, power and engineering at a scale only a few labs can afford.

Serving the result is its own bill: memory-bound decode, growing KV caches, long reasoning traces.

Next, deck 5: why we picked a research area where a student team can still reach the edge.

Elapsed
0:00:00
Steps on this slide
Next slide