QMIND Research · Lecture 6 of 7

Dataset engineering and experiment design

The guts of the project: how we manufacture the data, hold out combinations, and sweep the dials that turn a demo into a paper.

Part A · The data factory

We don't download a dataset. We manufacture one.

Downloaded
  • Factors tangled and unlabelled
  • Missing combinations stay missing
  • No answer key for a new image
Any combination on demand, each with an exact answer key. So we can ask what nobody with a scraped dataset can: what happens on combinations the model has never seen?

SCAN (Lake & Baroni, ICML 2018) and CLEVR (Johnson et al., CVPR 2017) were built synthetically for the same reason: control over composition.

Part A · The data factory

Every frame flows through one deterministic pipeline

The skinner is the answer key. It paints every training target and, later, grades every test frame, including combinations the model never saw.

One ray per screen column, 64 per frame, no GPU. Each worker owns its seeds, so generation is embarrassingly parallel. Cost: [ms/frame: measure in week 1].

Part A · The data factory

One training example is a stack of aligned tensors

Structure · input
Answer key · target

The U-Net sees : 15 channels. The theme enters as a global code.

Labels + metadata
theme=Lava · material=bricks · lighting=Dusk
map_seed=4127 · cam_step=37 · data=v0.1
Reproducible from five numbers
Part A · The data factory

Pixels live in shards; the truth lives in a manifest

manifest.parquet · one row per frame
FieldTypeExample
example_idstrm4127_k037_t1m0l1
map_seed, cam_stepint4127, 37
theme, material, lightingint1, 0, 1
splitenumtrain · val · test
shardstrshard-00012.tar
sha256hex9f2c…e1a0
gen_commitgit hasha1b2c3d
data_versionstrv0.1

About 53 KB per frame raw (13 bytes × 4,096 px), so 128k frames ≈ 7 GB. proposed

Part A · The data factory

Trust nothing you haven't plotted

$ pytest tests/data -qtest_same_seed_same_bytes      PASSEDtest_pattern_sticks_to_wall    PASSEDtest_theme_independent_of_map  PASSEDtest_no_test_cell_in_train     PASSED
Part A · The data factory

Factors are gridded; everything else is randomized

Trap: if Neon maps were always bigger, the model could read the theme off the geometry and "compose" by cheating.

Fix: draw every nuisance knob from the same distribution in every cell, then test it. A probe that predicts a factor from the structure channels alone should score at chance.

Part B · Held-out compositions

The factor grid: every cell is one combination

4 themes × 4 materials × 4 lightings = 64 cells · ~2,000 frames per cell (40 maps × 50 camera steps) ≈ 128k frames

Part B · Held-out compositions

Hold out combinations, never values

Hold out a value = extrapolation. The model has never seen Neon at all, so nothing tells it what Neon looks like.

Hold out a pairing = composition. Seen Neon. Seen Torch. Never together.

Same atoms, new compounds, as in compound-divergence splits (Keysers et al., 2020). Diffusion models compose seen factors far better than they interpolate to unseen values (Liang et al., 2024).

Keysers et al., "Measuring Compositional Generalization," ICLR 2020 · Liang, Liu, Ostrow, Fiete, "How Diffusion Models Learn to Factorize and Compose," NeurIPS 2024

Part B · Held-out compositions

Three holdout schemes, from easy to hard

Random cells. About 15% of cells, scattered. Every test cell has seen neighbours. Tests: filling holes in the grid.

Structured block. A whole slice: Jungle under Torch and Dark, every material. No close neighbours. Tests: a correlated gap.

Leave one theme out. Neon never appears. Only a parameter-vector encoding can try: an ID has no embedding for a theme it never saw.

In every scheme, validation cells are separate from test cells, and all of it is fixed before training and written into the manifest.

Part B · Leakage

Leakage: a test frame that is secretly a training frame

Audit: for every test frame, find its nearest training frame. Anything closer than a threshold is a leak, whatever the split code says.

Part B · Leakage

Four ways a test set quietly stops being a test

Trap 1 · shared maps

Same map in train and test

Split by map seed (and camera path), never by frame.

Trap 2 · near-duplicates

Consecutive camera steps

Group split, plus a nearest-neighbour audit.

Trap 3 · tuning on test

Choosing settings by test score

Tune on validation cells. Touch test cells once.

Trap 4 · label leaks

Factor readable from another channel

Sample factors independently; probe the structure channels.

Part B · Held-out compositions

Compare like with like: same maps, every cell

Same maps, camera paths, frame counts and noise seeds in every cell. Only the combination changes.

Train mapsNew maps
Seen cellin-distributionnew layout
Unseen cellnew combinationboth
Part C · Sweeps

One dial per sweep. The answer is a curve.

Held fixed: total training frames · U-Net · training steps · test cells · seeds

SweepThe dialTests
Coverage20 → 90% of cells seenH1 threshold
Encodingpixel map · ID · vectorH2 local vs global
Interactionper-pixel → shadows, reflectionsH3 non-local
Factor count3 → 6+ factorsH4 scale
Data, modelframes, U-Net widthcontrols
If it succeeds everywhere, keep turning until it fails. The breaking point is the result.
Part C · Sweep 1 · Coverage

Coverage: is there a threshold?

Threshold: H1 holds. Report where it sits, and whether it moves as factors grow.

Smooth: no phase transition. Report the slope.

Flat at chance: it memorizes cells. The paper becomes why.

Design: test cells fixed, training cells nested as coverage grows, total frames constant.

Multiplicative emergence: Okawa, Lubana, Dick, Tanaka, "Compositional Abilities Emerge Multiplicatively," NeurIPS 2023

Part C · Sweep 2 · Encoding

Encoding: how the theme reaches the network

Part C · Sweep 3 · Interaction

Interaction: local effects vs effects that span the scene

Why expect this: convolutional diffusion models behave like local patch mosaics, built from locality and equivariance (Kamb & Ganguli, 2025). Non-local effects are where that should break.

Kamb & Ganguli, "An analytic theory of creativity in convolutional diffusion models," ICML 2025

Part C · Sweep 4 + controls

More factors means rarer combinations

Controls: if more data or a wider U-Net closes the gap, the limit was capacity. If the gap persists, it's composition.

Rarer concepts need more optimisation before they compose: Okawa et al., NeurIPS 2023

Part C · Compute budget

The whole study fits in the AWS credits, on paper

SweepSettingsSeedsRuns
Coverage × 2 schemes6 × 2336
Encoding × 2 schemes3 × 2318
Interaction339
Factor count4312
Data & model size3 + 3318
Baselines, ablations5315
Total108
Assumptions (estimates)

1× A10G (g5.xlarge): $1.006/h on-demand, us-east-1

~6 GPU-h per run [our estimate: measure at Gate 2]

: every run paid for twice (pilots, bugs, reruns)

Price: AWS g5.xlarge on-demand, us-east-1, $1.006/h; spot was about $0.40/h (Holori AWS price calculator, Oct 2026). GPU-hours per run are our estimate.

Part D · Measuring

Every generated frame has an exact answer key

PSNR · pixel error
LPIPS · perceptual distance

PSNR rewards blurry averages; LPIPS punishes them. Report both. LPIPS was calibrated on 64×64 patches: our frame size.

LPIPS: Zhang, Isola, Efros, Shechtman, Wang, "The Unreasonable Effectiveness of Deep Features as a Perceptual Metric," CVPR 2018 (BAPPS uses 64×64 patches)

Part D · Measuring

Per-factor classifiers catch an ignored signal

Report generated-frame accuracy relative to the classifier's accuracy on skinner frames (the ceiling).

Part D · Measuring

The headline: the seen-vs-unseen gap, cell by cell

The pattern in the failures is the paper: which combinations break tells us why.

Part D · Secondary metrics

Does it flicker? Warp the last frame and compare

Frame time: ms per frame vs number of DDIM steps. Real time at 30 fps is 33 ms.

Part D · Baselines + ablations

Baselines bracket the result; ablations explain it

Ablations: remove one part, re-measure
  • Drop UV: do patterns still stick to walls?
  • Drop the light map: can it light from the factor alone?
  • Theme via adaptive norm vs concatenation
  • DDIM steps: 50 → 4
  • Fixed vs random starting noise
Part D · Rigor

At least three seeds, and bootstrap over cells

Resample cells, not frames: frames inside a cell are correlated, so treating them as independent fakes precision.

# every run, logged automaticallyrun     cov50-random-s2config  configs/coverage.yamlgit     3f9a1c2  data v0.3seed    2   → W&B or MLflow

Stratified bootstrap CIs over tasks: Agarwal et al., "Deep RL at the Edge of the Statistical Precipice," NeurIPS 2021

Part D · The paper's figures

Sketch the figures before the first run

Part E · The paper

What turns this into a paper

1

A clean question

Where does composition break, and why?

2

Controlled data

Every factor a dial; exact answer keys.

3

Curves with error bars

Sweeps, seeds, bootstrap CIs.

4

An explanation

Where it breaks, and a tested reason why.

5

Negative results count

"It fails here, for this reason" is a finding.

Our niche

Prior work probes composition in diffusion on toy shapes and 2-D Gaussians. Ours is a full rendering pipeline, with a client's real question.

Okawa et al., NeurIPS 2023 · Park et al., "Emergence of Hidden Capabilities," NeurIPS 2024 · Liang et al., NeurIPS 2024

Part E · The plan

Five phases, four gates

Each gate is make-or-break. Dates are set with the team: [gate dates TBD]. Gate 3 is the biggest risk; if it fails, the failure is what we explain.

Part E · The call

Five streams. Here's your first two weeks.

Stream 1

Engine & rendering

  • Ray-caster emits 5 channels at 64×64
  • Map from a seed
  • Determinism test

Need: geometry, NumPy

Stream 2

Skinning & data

  • Theme record + 4 patterns
  • skin() v0, contact sheets
  • Shard writer + manifest

Need: procedural art, plumbing

Stream 3

Model & training

  • U-Net + DDPM loop
  • Overfit one batch
  • First non-static sample

Need: PyTorch, patience

Stream 4

Evaluation & analysis

  • Split generator: 3 schemes + val
  • PSNR/LPIPS harness
  • Factor classifiers v0

Need: stats, skeptics

Stream 5

Infra & MLOps

  • AWS + S3 bucket
  • Training container
  • Tracking + cost alarms

Need: cloud, Docker, Linux

THE CALL

“The first held-out cell rendered correctly will be ours to explain.”

Pick a stream. Bring a laptop. We start building at the next meeting.

Elapsed
0:00:00
Steps on this slide
Next slide