QMIND Research · Lecture 5 of 7

Our research, and how diffusion works

Why we picked neural rendering of games, how a denoiser turns static into a frame, and exactly how our game gets into the network.

PART A · NOW, OUR PROJECT

Okay, that was cool. Now let's talk about our project.

Deck 2

One neuron learns

Deck 3

Transformers

Deck 4

The frontier at datacentre scale

So where, in all of that, can a student team do research that actually matters?

Why not LLMs

The LLM frontier is a capital race we can't enter

>$700B

planned 2026 capex: Alphabet, Amazon, Microsoft, Meta

~90%

of notable 2024 AI models came from industry

~5 mo

doubling time of training compute

So: pick a question where scale isn't the answer.

Sources: Sherwood News, Apr 30 2026 (2026 capex plans) · Stanford HAI, AI Index 2025

The map

Where can a student team do real research?

01

Interpretability of small models

Reverse-engineer what a small network computes.

02

Efficient fine-tuning

Adapt pretrained models cheaply: LoRA, quantization.

03

RL for small domains

Agents for games, puzzles and control tasks.

04

Robotics, sim-to-real

Train in simulation, run on real hardware.

05

Scientific ML

Learned surrogates for physics, chemistry, biology.

06

Generative models of structured worlds

Neural rendering, world models.

Scoring on deck 1's ROI factors

Scored on our constraints, one area stands out

Scores are Matt's judgement on a 1–5 scale for a student team with $10k, not measurements.

Second pass on deck 1

Why neural rendering of games won

  1. Small compute. 64×64 frames: 12,288 numbers each.
  2. Exact ground truth. The skinner renders the right answer for any combination.
  3. A real open question. Where does compositional generalization break?
  4. A paying client. Chatforce wants live re-theming.
  5. Close to hot work. GameNGen, DIAMOND, Genie.
Close to hot work

We sit next to the hottest work in generative AI

Alonso et al. NeurIPS 2024 · Valevski et al. ICLR 2025 · Decart/Etched · DeepMind Genie 2, 3, Project Genie · Microsoft Muse (Nature 2025)

The research question

The question: where does generalization break?

Hold out combinations, never values: every theme and material is seen, some pairings never are.

H1 · Coverage

Held-out accuracy jumps at a threshold of grid coverage.

H2 · Factor type

Pixel-aligned channels compose better than a global theme code.

H3 · Interaction

Non-local effects (shadows, reflections) generalize worse.

H4 · Scale

More factors, worse; rarest combinations fail first.

Make-or-break

Three moments where the project lives or dies

Moment 1

A denoised image that isn't static

Moment 2

A held-out combination rendered right

Moment 3

A curve showing where it breaks

If moment 2 fails, the failure becomes the finding: we explain why.

What you get out of it

Five streams, and what each one teaches you

StreamYou buildYou learn
Engine & renderingray-caster, map generator, channel outputsray casting, procedural generation, determinism
Skinning & datatheme space, skinner, dataset generatorprocedural textures, parameter spaces, data at scale
Model & trainingU-Net, conditioning, samplersdiffusion from scratch, architecture, training dynamics
Evaluation & analysisholdout splits, metrics, classifiersexperimental design, statistics, honest claims
Infrastructure & MLOpsAWS training jobs, storage, trackingcloud GPUs, pipelines, reproducibility

Everyone owns a piece of one stack: build the world, make the data, train the model, design the test.

The paper story

The paper we're aiming for, in three sentences

Setup

Neural renderers like GameNGen work, but nobody can say where they fail, because real games have no answer key.

Method

We build a game whose answer key is exact, and train on only some combinations of theme, material and lighting.

Result

A map of where compositional generalization breaks, and evidence for why.

CUCAI paperWorkshop submissionPlayable demoA negative result is still a result.
PART B

How diffusion works

Destroy a frame with noise, a little at a time. Then train a network to undo one little step.

Diffusion

Generative modelling means learning a distribution

Training data: samples from an unknown .

Learn , or something that lets you sample from it.

Sample: new points that look like the data but aren't copies.

A 64×64 frame is one point in . Real frames sit on a thin, curved surface: the data manifold.

Diffusion · forward process

Forward: add a little noise, a thousand times

: fresh Gaussian noise every step. grows from to over steps.

Shrink the image slightly, add a little noise.

By it is indistinguishable from pure noise.

Nothing here is learned. It's a fixed recipe.

Schedule: Ho, Jain & Abbeel, "Denoising Diffusion Probabilistic Models", NeurIPS 2020 (linear , T = 1000).

Diffusion · closed form

You can jump straight to any noise level

signal weight
noise weight
cosine schedule

Cosine schedule: Nichol & Dhariwal, "Improved Denoising Diffusion Probabilistic Models", ICML 2021.

Diffusion · the key trick

The trick: undo the noise one small step at a time

forward : fixed
reverse : learned
one giant jump: not learnable

For small , each reverse step is almost Gaussian, so the network only has to predict its mean.

Sohl-Dickstein et al., "Deep Unsupervised Learning using Nonequilibrium Thermodynamics", ICML 2015 · Ho et al., NeurIPS 2020.

Diffusion · the step function

The denoising step is one function

noisy frame
noise level
conditioning
predicted noise
clean guess
Diffusion · training

Training: add known noise, then predict it

1 sample a frame with its channels
2 pick at random, draw
3 jump:
4 predict
5 loss, backprop, update

Plain mean-squared error. The target is the noise we added, so we always know the right answer.

Diffusion · training

The whole training step fits on one screen

def train_step(model, x0, cond, theme):    B = x0.shape[0]    t = torch.randint(0, T, (B,))      # noise level    eps = torch.randn_like(x0)          # the answer    ab = alpha_bar[t].view(B, 1, 1, 1)    xt = ab.sqrt()*x0 + (1-ab).sqrt()*eps    inp = torch.cat([xt, cond], dim=1)  # 3+13=16    theme = drop(theme, p=0.1)       # for CFG    eps_hat = model(inp, t, theme)    loss = F.mse_loss(eps_hat, eps)    opt.zero_grad(); loss.backward()    opt.step()    return loss
Concatenate

Structural channels ride along with the noisy image. Part D.

Global inputs

and theme go in separately, as vectors.

Same as deck 2

Loss, backward, step. Only the data recipe is new.

Diffusion · sampling

Sampling: start from noise, step back to a frame

Start: . For :

Remove a bit of the predicted noise, then add a little fresh noise .

Original DDPM: 1000 network calls per image.

Diffusion · samplers

DDPM adds fresh noise each step; DDIM adds none

DDIM: same trained network, no fresh noise. Same start, same frame. Every time.

Skips steps safely: 10–50× faster.

For a game: same input, same frame, so less flicker.

Song, Meng & Ermon, "Denoising Diffusion Implicit Models", ICLR 2021.

Intuition 1

The predicted noise points toward the data

The score: an arrow toward higher probability.

At low noise the arrows hug the data; at high noise they only know roughly where it is.

Sampling = follow the arrows while turning the noise down.

Annealed Langevin dynamics: Song & Ermon, NeurIPS 2019 · score SDEs: Song et al., ICLR 2021.

Intuition 2

Intuition 2: coarse first, fine detail last

High noise decides layout; low noise decides texture. Diffusion is roughly autoregression over frequencies.

Dieleman, "Diffusion is spectral autoregression", blog post, 2024 · natural-image spectra fall off roughly as .

Intuition 3

Intuition 3: one big jump lands on the average

The MSE-best single guess is the average of every plausible answer: a blur. Small steps commit one detail at a time.

Intuition 4

Intuition 4: when does it just memorize?

Kadkhodaie et al., ICLR 2024

Faces: up to ~100 training images, samples are copies. At , two models trained on disjoint data generate nearly identical new faces.

Carlini et al., 2023

Over a thousand training images extracted from large text-to-image models.

Held-out combinations are our memorization test.

Kadkhodaie, Guth, Simoncelli & Mallat, ICLR 2024 · Carlini et al., "Extracting Training Data from Diffusion Models", USENIX Security 2023.

Conditioning & guidance

Conditioning, and turning it up with guidance

One network, trained with dropped 10–20% of the time, gives both terms.

: sharper and more on-theme, less diverse.

Cost: two network calls per step.

Ho & Salimans, "Classifier-Free Diffusion Guidance", NeurIPS 2021 Workshop (arXiv 2022).

Inside · the U-Net

The U-Net: shrink to understand, grow back to place

Ronneberger, Fischer & Brox, "U-Net: Convolutional Networks for Biomedical Image Segmentation", MICCAI 2015 · sizes shown are our planned config.

U-Net · building block

Every block: normalize, activate, convolve, add back

h GroupNorm SiLU Conv 3×3 GroupNorm scale, shift SiLU Conv 3×3 + from t + theme embedding residual: output = h + F(h)
GroupNorm

Normalize each image's channels in groups (e.g. 32). Doesn't depend on batch size, so small batches are fine.

SiLU

: a smooth ReLU.

Residual

The block learns a correction to . Gradients flow straight through the skip, as in deck 3's residual stream.

U-Net · the noise level

The timestep becomes a vector that reaches every block

U-Net · attention

Self-attention at 16×16 lets distant pixels agree

Each of the 256 cells is a token: , , exactly as in deck 3.

A wall cell attends to the same wall across the frame: one pattern, one lighting.

Quadratic cost is why attention lives only at the lowest resolutions.

U-Net · our planned config

Our planned U-Net: about 16.5 million parameters

LevelSizeChannelsBlocks
Input64×6416 → 641 conv
Level 164×64642 res
Level 232×321282 res
Level 316×162562 res + attn
Middle16×16256res, attn, res
Decodermirror256 → 643 res / level
Output64×6464 → 31 conv

≈17.5 GFLOPs per denoising step at 64×64.

MNIST net: 13,002DDPM CIFAR-10: 35.7MOurs: ~16.5MWidth 128: ~66M

Plan, not final: base width 64, multipliers (1, 2, 4), AdaGN conditioning; counts computed by us. DDPM figure: Ho et al. 2020.

How our inputs get in

Every pixel arrives as a stack of 16 numbers

How our inputs get in

Concatenation works because the first conv sees it all

Each output mixes all 16 channels around pixel .

SR3: low-res image concatenated SD inpainting: 9 input channels GameNGen: past frames in latent channels ControlNet: depth and edge maps
How our inputs get in

Global inputs turn per-channel dials

FiLM: Perez et al., AAAI 2018 · AdaGN: Dhariwal & Nichol, "Diffusion Models Beat GANs on Image Synthesis", NeurIPS 2021 · dial values illustrative.

Theme encoding · hypothesis 2

Theme as an ID can't generalize; as parameters, it might

Temporal stability

Three tricks keep frames from flickering

No autoregressive drift: the engine owns gameplay, so every frame is rendered from exact structure, not from our last guess.

Schematic: hallucinated detail exaggerated for visibility.

Inference

Few-step sampling makes a playable demo plausible

Cost per frame = steps × one U-Net call (≈17.5 GFLOPs).

At 20 fps with 4 steps: ≈1.4 TFLOP/s, a small fraction of one modern GPU's peak.

GameNGen: 4 DDIM steps, 20 fps on one TPU.

GameNGen figures: Valevski et al., 2024. FLOP counts are our estimates for the planned config; real frame time is a measurement we'll make.

Next: deck 6

Now: what do we train this on, and how do we prove it generalizes?

The model is the easy part to describe. The data and the tests are where the science is.

Elapsed
0:00:00
Steps on this slide
Next slide