QMIND Research · Lecture 5 of 7
Our research, and how diffusion works
Why we picked neural rendering of games, how a denoiser turns static into a frame, and exactly how our game gets into the network.
PART A · NOW, OUR PROJECT
Okay, that was cool. Now let's talk about our project.
Deck 4
The frontier at datacentre scale
So where, in all of that, can a student team do research that actually matters?
Why not LLMs
The LLM frontier is a capital race we can't enter
>$700B
planned 2026 capex: Alphabet, Amazon, Microsoft, Meta
~90%
of notable 2024 AI models came from industry
~5 mo
doubling time of training compute
So: pick a question where scale isn't the answer.
Sources: Sherwood News, Apr 30 2026 (2026 capex plans) · Stanford HAI, AI Index 2025
The map
Where can a student team do real research?
01
Interpretability of small models
Reverse-engineer what a small network computes.
02
Efficient fine-tuning
Adapt pretrained models cheaply: LoRA, quantization.
03
RL for small domains
Agents for games, puzzles and control tasks.
04
Robotics, sim-to-real
Train in simulation, run on real hardware.
05
Scientific ML
Learned surrogates for physics, chemistry, biology.
06
Generative models of structured worlds
Neural rendering, world models.
Scoring on deck 1's ROI factors
Scored on our constraints, one area stands out
Scores are Matt's judgement on a 1–5 scale for a student team with $10k, not measurements.
Second pass on deck 1
Why neural rendering of games won
- Small compute. 64×64 frames: 12,288 numbers each.
- Exact ground truth. The skinner renders the right answer for any combination.
- A real open question. Where does compositional generalization break?
- A paying client. Chatforce wants live re-theming.
- Close to hot work. GameNGen, DIAMOND, Genie.
Close to hot work
We sit next to the hottest work in generative AI
Alonso et al. NeurIPS 2024 · Valevski et al. ICLR 2025 · Decart/Etched · DeepMind Genie 2, 3, Project Genie · Microsoft Muse (Nature 2025)
The research question
The question: where does generalization break?
Hold out combinations, never values: every theme and material is seen, some pairings never are.
H1 · Coverage
Held-out accuracy jumps at a threshold of grid coverage.
H2 · Factor type
Pixel-aligned channels compose better than a global theme code.
H3 · Interaction
Non-local effects (shadows, reflections) generalize worse.
H4 · Scale
More factors, worse; rarest combinations fail first.
Make-or-break
Three moments where the project lives or dies
Moment 1
A denoised image that isn't static
Moment 2
A held-out combination rendered right
Moment 3
A curve showing where it breaks
If moment 2 fails, the failure becomes the finding: we explain why.
What you get out of it
Five streams, and what each one teaches you
| Stream | You build | You learn |
| Engine & rendering | ray-caster, map generator, channel outputs | ray casting, procedural generation, determinism |
| Skinning & data | theme space, skinner, dataset generator | procedural textures, parameter spaces, data at scale |
| Model & training | U-Net, conditioning, samplers | diffusion from scratch, architecture, training dynamics |
| Evaluation & analysis | holdout splits, metrics, classifiers | experimental design, statistics, honest claims |
| Infrastructure & MLOps | AWS training jobs, storage, tracking | cloud GPUs, pipelines, reproducibility |
Everyone owns a piece of one stack: build the world, make the data, train the model, design the test.
The paper story
The paper we're aiming for, in three sentences
Setup
Neural renderers like GameNGen work, but nobody can say where they fail, because real games have no answer key.
Method
We build a game whose answer key is exact, and train on only some combinations of theme, material and lighting.
Result
A map of where compositional generalization breaks, and evidence for why.
CUCAI paperWorkshop submissionPlayable demoA negative result is still a result.
PART B
How diffusion works
Destroy a frame with noise, a little at a time. Then train a network to undo one little step.
Diffusion
Generative modelling means learning a distribution
Training data: samples from an unknown .
Learn , or something that lets you sample from it.
Sample: new points that look like the data but aren't copies.
A 64×64 frame is one point in . Real frames sit on a thin, curved surface: the data manifold.
Diffusion · forward process
Forward: add a little noise, a thousand times
: fresh Gaussian noise every step. grows from to over steps.
Shrink the image slightly, add a little noise.
By it is indistinguishable from pure noise.
Nothing here is learned. It's a fixed recipe.
Schedule: Ho, Jain & Abbeel, "Denoising Diffusion Probabilistic Models", NeurIPS 2020 (linear , T = 1000).
Diffusion · closed form
You can jump straight to any noise level
Cosine schedule: Nichol & Dhariwal, "Improved Denoising Diffusion Probabilistic Models", ICML 2021.
Diffusion · the key trick
The trick: undo the noise one small step at a time
one giant jump: not learnable
For small , each reverse step is almost Gaussian, so the network only has to predict its mean.
Sohl-Dickstein et al., "Deep Unsupervised Learning using Nonequilibrium Thermodynamics", ICML 2015 · Ho et al., NeurIPS 2020.
Diffusion · the step function
The denoising step is one function
Diffusion · training
Training: add known noise, then predict it
1 sample a frame
with its channels
Plain mean-squared error. The target is the noise we added, so we always know the right answer.
Diffusion · training
The whole training step fits on one screen
def train_step(model, x0, cond, theme): B = x0.shape[0] t = torch.randint(0, T, (B,)) # noise level eps = torch.randn_like(x0) # the answer ab = alpha_bar[t].view(B, 1, 1, 1) xt = ab.sqrt()*x0 + (1-ab).sqrt()*eps inp = torch.cat([xt, cond], dim=1) # 3+13=16 theme = drop(theme, p=0.1) # for CFG eps_hat = model(inp, t, theme) loss = F.mse_loss(eps_hat, eps) opt.zero_grad(); loss.backward() opt.step() return loss
Concatenate
Structural channels ride along with the noisy image. Part D.
Global inputs
and theme go in separately, as vectors.
Same as deck 2
Loss, backward, step. Only the data recipe is new.
Diffusion · sampling
Sampling: start from noise, step back to a frame
Start: . For :
Remove a bit of the predicted noise, then add a little fresh noise .
Original DDPM: 1000 network calls per image.
Diffusion · samplers
DDPM adds fresh noise each step; DDIM adds none
DDIM: same trained network, no fresh noise. Same start, same frame. Every time.
Skips steps safely: 10–50× faster.
For a game: same input, same frame, so less flicker.
Song, Meng & Ermon, "Denoising Diffusion Implicit Models", ICLR 2021.
Intuition 1
The predicted noise points toward the data
The score: an arrow toward higher probability.
At low noise the arrows hug the data; at high noise they only know roughly where it is.
Sampling = follow the arrows while turning the noise down.
Annealed Langevin dynamics: Song & Ermon, NeurIPS 2019 · score SDEs: Song et al., ICLR 2021.
Intuition 2
Intuition 2: coarse first, fine detail last
High noise decides layout; low noise decides texture. Diffusion is roughly autoregression over frequencies.
Dieleman, "Diffusion is spectral autoregression", blog post, 2024 · natural-image spectra fall off roughly as .
Intuition 3
Intuition 3: one big jump lands on the average
The MSE-best single guess is the average of every plausible answer: a blur. Small steps commit one detail at a time.
Intuition 4
Intuition 4: when does it just memorize?
Kadkhodaie et al., ICLR 2024
Faces: up to ~100 training images, samples are copies. At , two models trained on disjoint data generate nearly identical new faces.
Carlini et al., 2023
Over a thousand training images extracted from large text-to-image models.
Held-out combinations are our memorization test.
Kadkhodaie, Guth, Simoncelli & Mallat, ICLR 2024 · Carlini et al., "Extracting Training Data from Diffusion Models", USENIX Security 2023.
Conditioning & guidance
Conditioning, and turning it up with guidance
One network, trained with dropped 10–20% of the time, gives both terms.
: sharper and more on-theme, less diverse.
Cost: two network calls per step.
Ho & Salimans, "Classifier-Free Diffusion Guidance", NeurIPS 2021 Workshop (arXiv 2022).
The U-Net: shrink to understand, grow back to place
Ronneberger, Fischer & Brox, "U-Net: Convolutional Networks for Biomedical Image Segmentation", MICCAI 2015 · sizes shown are our planned config.
U-Net · building block
Every block: normalize, activate, convolve, add back
GroupNorm
Normalize each image's channels in groups (e.g. 32). Doesn't depend on batch size, so small batches are fine.
Residual
The block learns a correction to . Gradients flow straight through the skip, as in deck 3's residual stream.
U-Net · the noise level
The timestep becomes a vector that reaches every block
U-Net · attention
Self-attention at 16×16 lets distant pixels agree
Each of the 256 cells is a token: , , exactly as in deck 3.
A wall cell attends to the same wall across the frame: one pattern, one lighting.
Quadratic cost is why attention lives only at the lowest resolutions.
U-Net · our planned config
Our planned U-Net: about 16.5 million parameters
| Level | Size | Channels | Blocks |
| Input | 64×64 | 16 → 64 | 1 conv |
| Level 1 | 64×64 | 64 | 2 res |
| Level 2 | 32×32 | 128 | 2 res |
| Level 3 | 16×16 | 256 | 2 res + attn |
| Middle | 16×16 | 256 | res, attn, res |
| Decoder | mirror | 256 → 64 | 3 res / level |
| Output | 64×64 | 64 → 3 | 1 conv |
≈17.5 GFLOPs per denoising step at 64×64.
MNIST net: 13,002DDPM CIFAR-10: 35.7MOurs: ~16.5MWidth 128: ~66M
Plan, not final: base width 64, multipliers (1, 2, 4), AdaGN conditioning; counts computed by us. DDPM figure: Ho et al. 2020.
How our inputs get in
Every pixel arrives as a stack of 16 numbers
How our inputs get in
Concatenation works because the first conv sees it all
Each output mixes all 16 channels around pixel .
SR3: low-res image concatenated
SD inpainting: 9 input channels
GameNGen: past frames in latent channels
ControlNet: depth and edge maps
How our inputs get in
Global inputs turn per-channel dials
FiLM: Perez et al., AAAI 2018 · AdaGN: Dhariwal & Nichol, "Diffusion Models Beat GANs on Image Synthesis", NeurIPS 2021 · dial values illustrative.
Theme encoding · hypothesis 2
Theme as an ID can't generalize; as parameters, it might
Temporal stability
Three tricks keep frames from flickering
No autoregressive drift: the engine owns gameplay, so every frame is rendered from exact structure, not from our last guess.
Schematic: hallucinated detail exaggerated for visibility.
Inference
Few-step sampling makes a playable demo plausible
Cost per frame = steps × one U-Net call (≈17.5 GFLOPs).
At 20 fps with 4 steps: ≈1.4 TFLOP/s, a small fraction of one modern GPU's peak.
GameNGen: 4 DDIM steps, 20 fps on one TPU.
GameNGen figures: Valevski et al., 2024. FLOP counts are our estimates for the planned config; real frame time is a measurement we'll make.
Next: deck 6
Now: what do we train this on, and how do we prove it generalizes?
The model is the easy part to describe. The data and the tests are where the science is.