QMIND Research · Lecture 6 of 7
Dataset engineering and experiment design
The guts of the project: how we manufacture the data, hold out combinations, and sweep the dials that turn a demo into a paper.
~0.5 min
Deck 5 gave us the model. This deck is the experiment: the data that trains it and the tests that grade it. Honestly, this is the part that makes it research.
Intuition: the U-Net is a commodity. What nobody else has is a dataset where every factor is a dial and every test has an exact answer key.
Background: each tile is a real frame from a toy version of our pipeline, written for these slides. The amber "?" tiles are held-out combinations.
Part A · The data factory
We don't download a dataset. We manufacture one.
Downloaded
Factors tangled and unlabelled
Missing combinations stay missing
No answer key for a new image
Any combination on demand, each with an exact answer key. So we can ask what nobody with a scraped dataset can: what happens on combinations the model has never seen?
SCAN (Lake & Baroni, ICML 2018) and CLEVR (Johnson et al., CVPR 2017) were built synthetically for the same reason: control over composition.
~1.5 min
Most ML projects start by downloading a dataset. We can't, and that's what makes the project work.
Scraped images: lighting is tangled up with location and style, nothing is labelled by factor, and if a combination isn't in the data you can't ask for it. There's also no ground truth for a new image.
Ours: every frame is a pure function of five numbers. Map seed, camera step, theme, material, lighting.
Turn the theme dial: same geometry, same camera, new paint.
Lighting dial: same scene under a torch.
Material dial: the wall pattern changes and stays glued to the walls.
That's the payoff. Because we manufacture data, we can render combinations on purpose, leave them out of training on purpose, and grade the model exactly.
Intuition: this is a controlled experiment, not field observation.
Analogy: a wind tunnel versus watching the weather. In the tunnel you set airspeed and angle yourself, one at a time.
If asked: "Isn't synthetic data a toy?" Simple, yes, but on purpose: control is what lets us answer a precise question. GameNGen (Valevski et al., 2024) simulated real Doom with diffusion, but it can't grade held-out combinations exactly. We can.
Part A · The data factory
Every frame flows through one deterministic pipeline
The skinner is the answer key. It paints every training target and, later, grades every test frame, including combinations the model never saw.
One ray per screen column, 64 per frame, no GPU. Each worker owns its seeds, so generation is embarrassingly parallel. Cost: [ms/frame: measure in week 1] .
~1.5 min
Here's the whole data factory, left to right. Nothing in it is random except through a seed.
A seed builds a map procedurally: walls, pillars and enemy positions on a 16×16 grid (our real maps can be bigger).
A camera path through the map. Camera step k is a position and heading along it.
The ray-caster renders channels, not pixels: depth, surface type, material ID, texture coordinates (UV), light. These are pixel-aligned.
The skinner, skin(channels, theme, lighting), paints the final colours. Theme and lighting are its knobs.
Out comes the ground-truth frame. The diffusion model's whole job is to learn to replace this box.
Live: the camera walks the path and every stage updates. A Doom-style ray-caster casts one ray per screen column, so a 64×64 frame needs only 64 rays. We'll measure the real cost per frame in week 1.
Intuition: the engine owns geometry, the skinner owns appearance, and the model only has to learn appearance.
If asked: "Why not render RGB directly in the engine?" Separating the channels from the skin is what lets us re-paint identical geometry under any theme, which is what makes held-out combinations possible.
Part A · The data factory
One training example is a stack of aligned tensors
The U-Net sees : 15 channels. The theme enters as a global code.
Labels + metadata
theme=Lava · material=bricks · lighting=Dusk map_seed=4127 · cam_step=37 · data=v0.1
Reproducible from five numbers
~1.5 min
Let's make "one example" concrete. Every sheet here is one 64×64 channel from the same frame.
Depth: one channel, a float per pixel.
Surface type, one-hot over ceiling, floor, wall, enemy: four sheets. Material ID, one-hot over 4 wall materials: three of those sheets are empty in this frame, which is exactly what one-hot means.
UV, two channels: where along the wall each pixel sits. Plus the light map. Total: 12 structural channels.
The target is the skinner's RGB frame. At training time the U-Net gets the noisy image concatenated with the 12 structure channels: 15 in, 3 out.
Every example also carries its factor labels and metadata. The labels are never fed to the model as pixels; they're for splitting and grading.
Determinism: given the map seed s, the camera step k and the three factors, the frame is byte-identical every time. That's a unit test, not a hope.
Intuition: the model gets a perfect blueprint (structure) and must learn the paint (appearance).
Watch out: the channel list is our proposal. The engine and skinning streams may add or merge channels in week 1. Keep the count in the config, not hard-coded.
If asked: "Why one-hot and not an integer?" An integer implies an order (material 3 is "more" than material 1). One-hot, or a learned embedding, doesn't.
Part A · The data factory
Pixels live in shards; the truth lives in a manifest
manifest.parquet · one row per frame
Field Type Example
example_id str m4127_k037_t1m0l1
map_seed, cam_step int 4127, 37
theme, material, lighting int 1, 0, 1
split enum train · val · test
shard str shard-00012.tar
sha256 hex 9f2c…e1a0
gen_commit git hash a1b2c3d
data_version str v0.1
About 53 KB per frame raw (13 bytes × 4,096 px), so 128k frames ≈ 7 GB. proposed
~1 min
Storage is boring until it breaks a result. Two pieces: shards for pixels, a manifest for facts.
Workers write frames into shards: a few thousand examples per file (WebDataset-style .tar, or zarr/.npz). Big sequential files are fast to stream from S3; millions of tiny PNGs are not.
The manifest is a table, one row per frame. Keys first: map seed and camera step.
Then the factor labels. This is the table every split is computed from.
Then bookkeeping: split, which shard, a content hash, the generator's git commit and a data version. If a frame changes, its hash changes, and we know.
Shards go to S3; the loader streams shuffled shards to the GPU. Size: depth, UV and light as float16, surface and material as bytes, RGB as bytes: 13 bytes a pixel, about 53 KB a frame, roughly 7 GB raw for 128k frames, before compression.
Intuition: shards are the warehouse, the manifest is the inventory system. You never search the warehouse; you query the inventory.
If asked: "Why not regenerate on the fly, since it's deterministic?" We can, for small experiments. Stored, versioned shards guarantee every run reads exactly the same bytes, and they decouple CPU generation from GPU training.
Part A · The data factory
Trust nothing you haven't plotted
$ pytest tests/data -q test_same_seed_same_bytes PASSED test_pattern_sticks_to_wall PASSED test_theme_independent_of_map PASSED test_no_test_cell_in_train PASSED
~1.5 min
Before a single GPU-hour: validate the data. Three tools, each catching a different class of bug.
Contact sheets: dozens of random frames on one image. Your eye is the best anomaly detector you have. Look at one every time the generator changes.
Coverage: count frames per cell. The zeros should sit exactly on the held-out cells, and every individual value (every theme, every lighting) should still have plenty of training frames.
Unit tests for the skinner and splitter: same seed gives same bytes; the pattern moves with the wall as the camera moves (UV anchoring); theme is statistically independent of map geometry; no test cell ever appears in a training shard.
And here's why the contact sheet matters: this tile is labelled Neon but painted with Lava's palette. An off-by-one in a theme lookup. Every metric would happily train on it.
Intuition: a silent data bug turns into a confident wrong result. Cheap checks up front beat a retraction later.
If asked: "Can't we just test against the manifest?" The manifest is written by the same code that might be buggy. The image is the ground truth, so we look at images.
Part A · The data factory
Factors are gridded; everything else is randomized
Trap: if Neon maps were always bigger, the model could read the theme off the geometry and "compose" by cheating.
Fix: draw every nuisance knob from the same distribution in every cell, then test it. A probe that predicts a factor from the structure channels alone should score at chance.
~1 min
Two kinds of variables in our data, and they get opposite treatment.
Factors under study: theme, material, lighting. Discrete values, enumerated cell by cell. These are what we hold out and measure.
Nuisance diversity: map complexity, room size, camera speed and turning, enemy count and pose, the diffusion noise seed. We randomize these so the model can't memorize one scene, but we never hold them out by cell.
The trap: a nuisance variable correlated with a factor leaks the factor. If one theme always comes with big rooms, "rendering Neon" might really be "rendering big rooms".
The fix is independence by construction, plus a test: train a tiny classifier to predict theme from the structure channels only. If it beats chance, something leaks.
Intuition: randomize what you don't study; control what you do.
Analogy: a drug trial randomizes age and diet across both arms so the only systematic difference is the drug.
If asked: "Could a nuisance knob become a factor later?" Yes, that's exactly how the factor-count sweep grows the grid: promote fog density or enemy type from nuisance to factor.
Part B · Held-out compositions
The factor grid: every cell is one combination
4 themes × 4 materials × 4 lightings = 64 cells · ~2,000 frames per cell (40 maps × 50 camera steps) ≈ 128k frames
~1 min
Now the core object of the whole study: the factor grid.
Each small cube is a cell: one combination of theme, material and lighting.
Three axes, four values each to start. The proposal says 4–5 values; we start small and grow only once the small grid works.
One cell, for example Neon × bricks × Torch, means every frame of every map rendered with that combination.
The numbers: 64 cells, about 2,000 frames each from 40 maps and 50 camera steps. Around 128k frames, which we just saw is a few GB.
For the rest of the talk I'll flatten the cube into four slices, one per material, each a theme × lighting grid. Same 64 cells.
Intuition: the grid turns "does it generalize?" into a map you can colour in, cell by cell.
Subtle point: a frame shows walls, floor and ceiling. We define "material" as the wall material of the scene, with floor and ceiling fixed per theme, so every frame belongs to exactly one cell. Decide this explicitly in week 1.
Part B · Held-out compositions
Hold out combinations, never values
Hold out a value = extrapolation. The model has never seen Neon at all, so nothing tells it what Neon looks like.
Hold out a pairing = composition. Seen Neon. Seen Torch. Never together .
Same atoms, new compounds, as in compound-divergence splits (Keysers et al., 2020). Diffusion models compose seen factors far better than they interpolate to unseen values (Liang et al., 2024).
Keysers et al., "Measuring Compositional Generalization," ICLR 2020 · Liang, Liu, Ostrow, Fiete, "How Diffusion Models Learn to Factorize and Compose," NeurIPS 2024
~1.5 min
The single most important design rule of the project. Here's one slice: themes down, lightings across.
Option one: hold out a value, the whole Neon row. That's an unfair test of a different skill. The model can't know Neon's palette if it never saw Neon. That's extrapolation.
Option two: hold out pairings. Neon appears in training (under other lights), Torch appears (with other themes), but Neon-under-Torch never does. The green outlines show where the model did see each ingredient.
Formally: every marginal is covered, some joints are zero. That's the definition of a compositional split.
This lines up with the literature. Google's CFQ work calls it "same atoms, different compounds" and found accuracy falls as compound divergence rises. Liang et al. found diffusion models compose seen factors well but interpolate to unseen values poorly. So composition is the interesting, answerable question.
Intuition: we test whether the model learned the factors separately, not whether it can guess an unseen colour.
Analogy: you've had chocolate and you've had chili, never chocolate-chili. Could you predict the taste? That's composition. Predicting a spice you've never tasted is extrapolation.
If asked: "Isn't leave-one-theme-out a value holdout?" With a parameter-vector encoding, no: the new theme's ingredients (palette, pattern, fog) each appear in other themes. It's composition one level down. More on the next slide.
Part B · Held-out compositions
Three holdout schemes, from easy to hard
Random cells. About 15% of cells, scattered. Every test cell has seen neighbours. Tests: filling holes in the grid.
Structured block. A whole slice: Jungle under Torch and Dark, every material. No close neighbours. Tests: a correlated gap.
Leave one theme out. Neon never appears. Only a parameter-vector encoding can try: an ID has no embedding for a theme it never saw.
In every scheme, validation cells are separate from test cells , and all of it is fixed before training and written into the manifest.
~1.5 min
All 64 cells, flattened into four material slices. Blue is train. We'll colour in three different splits.
Random cells: hold out 10–20% scattered at random (here 10 test, 5 validation). Each test cell has seen cells on every side. Easiest: it's interpolation inside the grid.
Structured block: hold out a whole correlated chunk, like every Jungle scene under Torch or Dark lighting. Now the nearest seen examples are further away. Harder.
Leave one theme out: Neon never appears in training at all. An ID encoding is structurally unable to do this (an unseen ID has an untrained embedding). A parameter vector might, if Neon's recipe is made of ingredients seen in other themes.
Validation cells are what we tune hyperparameters and early stopping on. Test cells are touched once, at the end. For leave-one-theme-out, validation is a different held-out theme. Because themes can be sampled from the parameter space, we can hold out several.
Intuition: each scheme moves the test further from anything seen, so together they map how far composition stretches.
If asked: "Why not just random cells?" Random cells are the friendliest case. A paper that only reports them overstates generalization. The schemes are a difficulty axis in their own right.
Part B · Leakage
Leakage: a test frame that is secretly a training frame
Audit: for every test frame, find its nearest training frame. Anything closer than a threshold is a leak, whatever the split code says.
~1 min
Here's the most common way to fool yourself. Ten consecutive camera steps from one map.
Naive split: assign each frame to train or test at random, the way you would shuffle any dataset.
Zoom in: test frame k=65 and training frame k=64 are almost identical; the difference map is nearly black. These numbers are computed live from the toy renders. The "test" score here is really memorization.
The fix is a group split: split by map seed. Every frame of map 4127 goes to train; every frame of map 9120 goes to test. No geometry is shared.
And audit it anyway: nearest-neighbour distance from each test frame to the training set. Code has bugs; distances don't lie.
Intuition: neighbouring video frames are nearly the same data point. Randomly splitting them is like putting the answer to question 5 in question 6.
If asked: "Held-out cells are unseen by construction, so why care?" Two reasons. The seen-cell test frames we compare against would leak. And for unseen cells, a shared map lets the model copy geometry-specific details it memorized from other cells.
Part B · Leakage
Four ways a test set quietly stops being a test
Trap 1 · shared maps
Same map in train and test Split by map seed (and camera path), never by frame.
Trap 2 · near-duplicates
Consecutive camera steps Group split, plus a nearest-neighbour audit.
Trap 3 · tuning on test
Choosing settings by test score Tune on validation cells. Touch test cells once.
Trap 4 · label leaks
Factor readable from another channel Sample factors independently; probe the structure channels.
~1 min
A checklist we'll put in the repo and run before every result.
Shared maps: the same geometry on both sides. Split by map seed.
Near-duplicates: consecutive frames. Group split, then audit with nearest-neighbour distances.
Tuning on test: if you pick the learning rate, the checkpoint or the guidance scale by looking at test cells, you've trained on them through your own choices. Use validation cells.
Label leaks: the held-out factor is readable from something else. Examples: material IDs numbered per theme instead of globally; a theme correlated with room size; fog density implying the theme.
Trap 4 is the sneakiest because the model succeeds and the success is fake. The probe from the previous part (predict the factor from structure channels alone) is our alarm.
Intuition: every one of these makes the model look better than it is. A suspiciously good result is a bug until proven otherwise.
If asked: "How do we stop ourselves from peeking at test?" Process, not willpower. The evaluation code reads test cells only with a flag, and final runs are logged with that flag set. Deck 7 covers the tracking.
Part B · Held-out compositions
Compare like with like: same maps, every cell
Same maps, camera paths, frame counts and noise seeds in every cell. Only the combination changes.
Train maps New maps
Seen cell in-distribution new layout
Unseen cell new combination both
~1 min
To compare held-out cells with seen cells fairly, they must differ in one thing only: the combination.
So the evaluation set is one map, one camera path, rendered under every cell. Here: the same frame across all 16 theme × lighting cells.
Blue borders are seen cells, amber are unseen. Same pixels of geometry, same frame count.
The headline number, the generalization gap, is mean error on unseen cells minus mean error on seen cells, measured on identical frames. Also fix the sampler's starting noise per frame so noise doesn't differ between cells.
Going further, a 2×2: train maps vs new maps, seen vs unseen cell. That separates "new combination" from "new layout". Rendering training maps under an unseen cell is safe: those exact frames never appeared in training because the cell is held out.
Intuition: a controlled comparison changes exactly one thing.
If asked: "Why does the 2×2 matter?" If unseen cells only fail on new maps, the problem is layout generalization, not composition. The 2×2 tells us which story we're in.
Part C · Sweeps
One dial per sweep. The answer is a curve.
Held fixed: total training frames · U-Net · training steps · test cells · seeds
Sweep The dial Tests
Coverage 20 → 90% of cells seen H1 threshold
Encoding pixel map · ID · vector H2 local vs global
Interaction per-pixel → shadows, reflections H3 non-local
Factor count 3 → 6+ factors H4 scale
Data, model frames, U-Net width controls
If it succeeds everywhere, keep turning until it fails. The breaking point is the result.
~1.5 min
How do we generate research rather than a demo?
The demo question is yes/no: does it generalize? A yes or no is one bit, and either answer is a weak paper.
The research question is where it breaks. Turn one dial and plot a curve. The shape of the curve is the finding.
One dial at a time means everything else held fixed. In particular total training frames: if more coverage also meant more data, we couldn't tell which helped.
Five sweeps, one per hypothesis, plus data and model size as controls. Next slides take them one at a time.
And the rule from the proposal: if the model succeeds everywhere, make the task harder until it fails. A curve that never bends has no story; the knee is the story.
Intuition: a curve contains every yes/no answer at once, plus where the answer flips.
Analogy: materials testing. You don't ask "is this beam strong?"; you load it until it fails and report the load.
If asked: "Why not sweep two dials together?" Later, for interactions. First establish each main effect cleanly; a 2-D sweep costs settings × settings runs.
Part C · Sweep 1 · Coverage
Coverage: is there a threshold?
Threshold: H1 holds. Report where it sits, and whether it moves as factors grow.
Smooth: no phase transition. Report the slope.
Flat at chance: it memorizes cells. The paper becomes why .
Design: test cells fixed, training cells nested as coverage grows, total frames constant.
Multiplicative emergence: Okawa, Lubana, Dick, Tanaka, "Compositional Abilities Emerge Multiplicatively," NeurIPS 2023
~1.5 min
Sweep 1. x: what fraction of the 64 cells training covers, 20% to 90%. y: per-factor classifier accuracy on the held-out cells (chance is 25% with four values). Seen cells sit near the top. Everything on this slide is hypothetical: we have no data yet.
Our hypothesis H1 is the amber shape: little generalization, then a sharp rise past some coverage.
Why expect a threshold? Okawa et al. showed that if getting a combination right needs every component skill, success is roughly a product of the parts. A product of smooth curves is a sharp curve. They saw sudden emergence in diffusion models on a synthetic task.
Alternative: smooth. Every extra cell buys a bit more. Still a result, and we report the slope.
Alternative: flat. It never composes, it memorizes cells. Then the paper is about why, and our baselines and probes carry it.
Every outcome is publishable because each one means something different. Design detail: training sets are nested (each coverage level adds cells to the last), test cells stay fixed, and total frames stay constant.
Intuition: we commit to what each shape would mean before we see it. That's what separates a test from a story told afterwards.
If asked: "Why hold total frames constant?" Otherwise 90% coverage also means more data, and we'd be measuring data size, not coverage. The data-size control sweep checks that separately.
Part C · Sweep 2 · Encoding
Encoding: how the theme reaches the network
~1 min
Sweep 2 changes one thing: how the theme is handed to the U-Net. The data and everything else stay the same.
Pixel map: paint the theme's parameters into extra per-pixel channels, concatenated with the input. Pixel-aligned, like material and light.
Global ID: theme number 3 looks up a learned vector, injected through adaptive group norm in every block, just like the timestep embedding from deck 5.
Parameter vector: the theme's recipe (palette, pattern, scale, finish, fog) goes through a small MLP into the same adaptive norm.
H2 predicts the pixel map composes most easily: the network combines local signals naturally. Dashed bars are what H2 would look like, not data.
Leave-one-theme-out separates them sharply. The ID simply can't: an unseen theme's embedding was never trained. The vector and pixel map can at least try. If the vector works there, we have shown a re-theme can generalize to new themes, which is Chatforce's practical question.
Intuition: an ID is a name; a parameter vector is a description. You can't render a name you've never heard; you might render a description.
If asked: "Isn't the pixel map wasteful, a constant over the image?" Yes, but that's what makes it a clean test of where information enters. If it wins, location of conditioning matters, not just content.
Part C · Sweep 3 · Interaction
Interaction: local effects vs effects that span the scene
Why expect this: convolutional diffusion models behave like local patch mosaics , built from locality and equivariance (Kamb & Ganguli, 2025). Non-local effects are where that should break.
Kamb & Ganguli, "An analytic theory of creativity in convolutional diffusion models," ICML 2025
~1 min
Sweep 3 changes the skinner itself: how far does information travel to set a pixel's colour?
Per-pixel: colour depends only on this pixel's own channels (material, UV, light). The amber pixel needs nothing else.
Shadows: whether this floor pixel is dark depends on an occluder somewhere else in the scene.
Reflections on a glossy floor: the pixel depends on the wall above it.
Wall-spanning patterns, like a mural: the pixel depends on its position along the whole wall.
H3: the further the dependency reaches, the worse composition gets. Kamb and Ganguli showed that locality and equivariance alone explain a lot of what convolutional diffusion models produce: patch mosaics. That predicts trouble exactly when a pixel needs distant context. The bars are hypothetical.
Intuition: composing local rules is easy because each patch can be solved on its own. Global rules need the whole scene.
If asked: "Doesn't the U-Net have attention at low resolution?" Yes, so it can carry some global context. Whether that's enough for held-out combinations is exactly what we measure.
Part C · Sweep 4 + controls
More factors means rarer combinations
Controls: if more data or a wider U-Net closes the gap, the limit was capacity. If the gap persists, it's composition.
Rarer concepts need more optimisation before they compose: Okawa et al., NeurIPS 2023
~1 min
Sweep 4: add factors. Candidates to promote from nuisance: fog density, enemy type, floor material, camera height.
Three factors with four values: 64 cells.
Six factors: 4,096 cells. Exponential.
At a fixed budget of 128k frames, frames per cell fall from 2,000 to about 31. Every combination becomes rare.
H4: generalization degrades as factors grow, and fails first on the rarest combinations. Okawa et al. found that concepts that are rarer in training need many more optimisation steps before the model composes them.
The controls are what make this convincing. Sweep data size and U-Net width. If the gap closes with scale, it was capacity. If it stays, it's a compositional limit. Both curves are hypothetical.
Intuition: the grid explodes, the data doesn't. Composition is how a model survives that gap.
If asked: "Couldn't we just generate more frames?" We can, generation is cheap, and that's the data-size control. The question is whether the model needs every combination's frames or can get by with the factors.
Part C · Compute budget
The whole study fits in the AWS credits, on paper
Sweep Settings Seeds Runs
Coverage × 2 schemes 6 × 2 3 36
Encoding × 2 schemes 3 × 2 3 18
Interaction 3 3 9
Factor count 4 3 12
Data & model size 3 + 3 3 18
Baselines, ablations 5 3 15
Total 108
Assumptions (estimates)
1× A10G (g5.xlarge): $1.006/h on-demand, us-east-1
~6 GPU-h per run [our estimate: measure at Gate 2]
: every run paid for twice (pilots, bugs, reruns)
Price: AWS g5.xlarge on-demand, us-east-1, $1.006/h; spot was about $0.40/h (Holori AWS price calculator, Oct 2026). GPU-hours per run are our estimate.
~1.5 min
Can we afford this? Runs = sweeps × settings × seeds.
Coverage: 6 levels × 2 schemes × 3 seeds = 36. Encoding 18, interaction 9, factor count 12, data and model size 18, baselines and ablations 15.
108 runs. Assumptions, stated loudly: one A10G at $1.006/h; about 6 GPU-hours per run, our guess for a small U-Net at 64×64 for roughly 100k steps. We measure the real number at Gate 2 and redo this slide.
Training: 108 × 6 × $1.006 ≈ $650.
Double it for pilots, bugs and reruns, add about 10% for evaluation sampling: roughly $1.4k.
That leaves about $3.6k of headroom for bigger grids, wider models, or the RL stretch goal. Spot instances can cut cost a lot if our jobs checkpoint and resume (deck 7).
Intuition: at 64×64, compute isn't the bottleneck; disciplined experiment design is.
Watch out: if a run takes 30 GPU-hours instead of 6, this becomes $7k and we must cut seeds or settings. That's why Gate 2 measures it before any sweep starts.
If asked: "Storage cost?" A few GB in S3 is cents a month. Compute dominates.
Part D · Measuring
Every generated frame has an exact answer key
LPIPS · perceptual distance
PSNR rewards blurry averages; LPIPS punishes them. Report both. LPIPS was calibrated on 64×64 patches: our frame size.
LPIPS: Zhang, Isola, Efros, Shechtman, Wang, "The Unreasonable Effectiveness of Deep Features as a Perceptual Metric," CVPR 2018 (BAPPS uses 64×64 patches)
~1.5 min
Part D: measuring. Our unfair advantage: for every generated frame, the skinner gives the exact right answer for the same channels and factors.
Left: the skinner's frame, the answer key. Middle: an illustrative "model output" (a degraded copy, not a real result).
PSNR: log of peak signal over mean squared error, in dB. Higher is better; identical images give infinity. The number shown is computed from this illustration.
LPIPS: pass both images through a pretrained network, unit-normalize the features at each layer, weight the channels with learned weights , take squared distances, average over space, sum over layers. Lower is better. It tracks human judgements of similarity.
Why both: a blurry average of all plausible outputs minimizes squared error, so PSNR can look fine on blur. LPIPS sees the blur. Nice coincidence: LPIPS was trained on 64×64 patches.
Intuition: PSNR asks "same pixels?", LPIPS asks "would a person see a difference?"
Watch out: compute metrics in one fixed colour space and range ([0,1] vs [-1,1] for LPIPS input), and always on the same frames per cell.
If asked: "Why not FID?" FID compares distributions and needs thousands of samples. We have paired ground truth per frame, which is a much stronger signal.
Part D · Measuring
Per-factor classifiers catch an ignored signal
Report generated-frame accuracy relative to the classifier's accuracy on skinner frames (the ceiling).
~1.5 min
PSNR tells you how wrong; it doesn't tell you what's wrong. This is the diagnostic metric.
Train a small CNN on skinner frames only, with three heads: which theme, which material, which lighting.
On a skinner test frame it reads all three correctly. Its accuracy on skinner frames is the ceiling.
Now feed it the model's output for a held-out cell, Neon × bricks × Torch. Theme right, material right, lighting reads "Noon". The model ignored the lighting signal and snapped to a combination it had seen. (This frame is an illustration.)
Over all held-out cells, a confusion matrix shows the pattern: mass moving off the diagonal towards seen pairings. The matrix here is illustrative.
Report it relative to the ceiling, so a weak classifier can't masquerade as a weak model.
Intuition: PSNR is a thermometer; the classifiers are a diagnosis.
Watch out: the classifier must not see generated frames in training, and it must be robust to small blur, or it will penalize harmless softness. Augment it with mild blur and noise.
If asked: "Couldn't the model fool the classifier?" Only if we trained against it. We don't; it's a fixed measuring instrument.
Part D · Measuring
The headline: the seen-vs-unseen gap, cell by cell
The pattern in the failures is the paper: which combinations break tells us why .
~1 min
The figure we most want: per-cell error over the whole grid, random-cells scheme. All values here are made up to show the figure, not results.
Amber outlines: test cells. Violet: validation cells.
Seen cells fill in: low error, measured on new maps.
Unseen cells: higher, and not uniformly.
Collapse to the headline: mean unseen minus mean seen, with a confidence interval. One number for the abstract.
But the structure is the real finding. If failures cluster where two extreme values meet (say, darkest lighting with the most saturated theme), that suggests an explanation we can then test.
Intuition: the average tells you how much it breaks; the map tells you where, and "where" leads to "why".
If asked: "Which metric is on the colour scale?" Any per-cell metric works: LPIPS, 1 − classifier accuracy, or the gap per cell. The paper shows LPIPS and classifier accuracy side by side.
Part D · Secondary metrics
Does it flicker? Warp the last frame and compare
Frame time: ms per frame vs number of DDIM steps. Real time at 30 fps is 33 ms.
~1 min
Two secondary metrics, for playability: temporal stability and speed.
Frames t−1 and t, a few camera steps apart.
We know depth and the exact camera motion, so we can reproject frame t−1 into frame t's view. Violet pixels are disocclusions (newly visible) and get masked out. The formula averages the squared difference over the visible mask M.
For the skinner, the warped frame matches almost perfectly: that's the floor, the warp error a perfectly stable renderer would get. This warp is computed live in the toy engine.
An illustrative flickery output (textures swim, fresh noise every frame) lights up the difference. We report its warp error as a ratio to the skinner's.
Frame time: milliseconds per frame against the number of DDIM sampling steps. Real-time needs about 33 ms at 30 fps.
Intuition: a stable renderer paints the same wall the same way when you move.
If asked: "How do we reduce flicker?" Condition on UV (patterns anchored to walls), fix the starting noise, use a deterministic DDIM sampler. Temporal conditioning on the warped previous frame is a stretch goal.
Part D · Baselines + ablations
Baselines bracket the result; ablations explain it
Ablations: remove one part, re-measure
Drop UV: do patterns still stick to walls?
Drop the light map: can it light from the factor alone?
Theme via adaptive norm vs concatenation
DDIM steps: 50 → 4
Fixed vs random starting noise
~1.5 min
A number means nothing without reference points. Baselines are the floor and ceiling around ours.
Ceiling: the skinner itself. Perfect by definition. PSNR is infinite, so it matters for the classifier and flicker metrics.
Floor: nearest-neighbour retrieval. Return the most similar training frame. It's the pure memorizer: if we can't beat it on unseen cells, we haven't learned to compose.
Two in between. "Nearest seen combination": the skinner rendered with the closest seen pairing, i.e. what ignoring the new factor looks like. And a regression U-Net: same architecture, predicts RGB directly with MSE, no diffusion. It tells us whether diffusion itself matters; expect it to be blurry.
Ours sits somewhere in the bracket. The placement shown is a guess, not a result.
Ablations remove one ingredient at a time to show which part does the work. Each one is a falsifiable claim about the mechanism.
Intuition: baselines answer "is it good?", ablations answer "why is it good?"
If asked: "Isn't the regression U-Net a strawman?" No. For a deterministic target it can be strong on PSNR. If it matches diffusion on held-out cells, that's a real finding about when generative modelling is needed.
Part D · Rigor
At least three seeds, and bootstrap over cells
Resample cells , not frames: frames inside a cell are correlated, so treating them as independent fakes precision.
# every run, logged automatically run cov50-random-s2 config configs/coverage.yaml git 3f9a1c2 data v0.3 seed 2 → W&B or MLflow
Stratified bootstrap CIs over tasks: Agarwal et al., "Deep RL at the Edge of the Statistical Precipice," NeurIPS 2021
~1.5 min
Two runs with different random seeds can disagree more than two methods do. So: seeds and error bars, always.
One run is one draw. Is that bump at 50% coverage real?
Three seeds (initialization, data order, and for random cells, which cells get held out) show the spread. Here the single run was a lucky draw. Illustrative data.
Bootstrap: resample the cells with replacement, recompute the mean, repeat B times (say 10,000). The middle 95% of those means is the confidence interval. No normality assumption needed.
Key subtlety: resample cells, not frames. The 2,000 frames in a cell are highly correlated, so treating them as independent makes intervals far too narrow. Agarwal et al. made the same point for RL: bootstrap over the units you generalize across.
And every run logs config, git hash, data version and seed automatically. If a number can't be traced to those four things, it doesn't go in the paper.
Intuition: error bars aren't decoration; they say which differences you're allowed to believe.
If asked: "Why not more seeds?" Budget. Three is the floor; we add seeds where the curves are close or near the knee.
Part D · The paper's figures
Sketch the figures before the first run
~1 min
A trick from good labs: draw the paper's figures, with empty axes, before running anything. If you can't sketch the figure, you don't yet know what experiment to run.
Fig 1, the system: engine, channels, skinner, renderer. Fig 2, per-cell error heatmaps with held-out cells outlined.
Fig 3, coverage curves with confidence bands, one line per holdout scheme. Fig 4, encoding comparison: pixel map, ID, vector across random, block and leave-one-theme-out.
Fig 5, the qualitative grid: rows are held-out cells, columns are skinner, ours, nearest-neighbour, regression. Reviewers look at this first. Table 1: main numbers with CIs.
Every figure maps to a hypothesis. The evaluation stream owns the scripts that make these straight from the tracking logs.
Intuition: the figure plan is the experiment plan, drawn.
If asked: "What if results don't fit the sketch?" Then the sketch was a hypothesis and the result is the news. Axes stay the same; the curves change.
Part E · The paper
What turns this into a paper
1
A clean question Where does composition break, and why?
2
Controlled data Every factor a dial; exact answer keys.
3
Curves with error bars Sweeps, seeds, bootstrap CIs.
4
An explanation Where it breaks, and a tested reason why.
5
Negative results count "It fails here, for this reason" is a finding.
Our niche
Prior work probes composition in diffusion on toy shapes and 2-D Gaussians. Ours is a full rendering pipeline, with a client's real question.
Okawa et al., NeurIPS 2023 · Park et al., "Emergence of Hidden Capabilities," NeurIPS 2024 · Liang et al., NeurIPS 2024
~1.5 min
Part E. Let's be explicit about what makes this publishable at CUCAI and a workshop.
A clean question, one sentence long, that the data can answer.
Controlled data, which is everything in Part A.
Curves with error bars, which is Parts C and D.
An explanation: not just where it breaks but a mechanism, tested with an ablation or a probe.
Negative results count. If the renderer never composes lighting, a clean demonstration of that, with a reason, is a result. That's why there's no way for this project to "fail" if we do the science properly.
Our niche: Okawa et al. and Park et al. studied composition in diffusion on synthetic shapes; Liang et al. on 2-D Gaussian blobs. We bring the same question to a structured rendering task with pixel-aligned conditioning, frames over time, and an industry client.
Intuition: reviewers reward a sharp question answered carefully more than a flashy demo.
If asked: "Do we need state-of-the-art numbers?" No. This is an analysis paper. The contribution is understanding, not a leaderboard.
Part E · The plan
Five phases, four gates
Each gate is make-or-break. Dates are set with the team: [gate dates TBD] . Gate 3 is the biggest risk; if it fails, the failure is what we explain.
~1 min
The plan in five phases. Each phase ends at a gate: a real test we either pass or learn from.
Phase 1, build the world. Gate 1: dataset v0 is deterministic, splits are clean, contact sheets look right.
Phase 2, learn to render. Gate 2: the first denoised frame that isn't static, then seen cells matching the skinner. Also where we measure GPU-hours per run.
Phase 3, held-out renders. Gate 3: the first held-out cell rendered right, or a failure we can explain. The biggest risk, and either outcome is content.
Phase 4, sweeps. Gate 4: the first curve with error bars showing where it breaks.
Phase 5: write it up for CUCAI and a workshop, and build a demo. Dates get set together at the next meeting.
Intuition: gates turn a year-long project into four checkpoints you can actually feel.
If asked: "What if we miss a gate?" We cut scope (smaller grid, fewer seeds), not rigor.
Part E · The call
Five streams. Here's your first two weeks.
Stream 1
Engine & rendering Ray-caster emits 5 channels at 64×64 Map from a seed Determinism test Need: geometry, NumPy
Stream 2
Skinning & data Theme record + 4 patterns skin() v0, contact sheetsShard writer + manifest Need: procedural art, plumbing
Stream 3
Model & training U-Net + DDPM loop Overfit one batch First non-static sample Need: PyTorch, patience
Stream 4
Evaluation & analysis Split generator: 3 schemes + val PSNR/LPIPS harness Factor classifiers v0 Need: stats, skeptics
Stream 5
Infra & MLOps AWS + S3 bucket Training container Tracking + cost alarms Need: cloud, Docker, Linux
~2 min
Here's where you come in. Five streams. Pick one tonight.
Engine: get the ray-caster emitting all five channels from a seed, and write the determinism test first.
Skinning and data: the theme record, four procedural patterns, skin() v0 with contact sheets, then the shard writer and manifest.
Model: a U-Net and DDPM loop on the full grid. First milestone is overfitting a single batch, the classic sanity check. Then the first non-static sample.
Evaluation: the split generator (all three schemes plus validation), the metric harness, and v0 factor classifiers. We need sceptical people here.
Infra: AWS account and S3 bucket, a training container, experiment tracking, and cost alarms before anyone launches a GPU.
The arrows show how they connect: each stream's output is the next one's input, and infra sits under all of them. Week-two integration target: one frame goes end to end.
Intuition: five small, testable deliverables beat one big integration at the end.
If asked: "Can I switch streams later?" Yes. Pick by what you want to learn, not what you already know.
THE CALL
“The first held-out cell rendered correctly will be ours to explain.”
Pick a stream. Bring a laptop. We start building at the next meeting.