noise⇄rendered frame
QMIND Research × Chatforce · 2026–27

Diffusion-Model Rendering for Re-Themable Games

Why this is worth a year of your time, and what we are going to build.

Return on effort
Conceptual illustration · not data

The domain you pick sets your return more than the hours do

Same input. Very different output.

The difference is where you point it.

Return on effort

Think of it as return on effort

D · Demand

Worth it now?

How much the skill is valued right now.

F · Frontier

How close?

Hours until you can do something new.

C · Compounding

Does it transfer?

Reused in the next thing you learn.

S · Scarcity

How rare?

How few people have the combination.

P · Proof

Can others see it?

Do you leave with evidence.

Return on effort
Conceptual sketch · not data

Every domain has a curve, and you want the steep part

API wrapper: a quick win, then flat.

Mature field: the easy wins are taken.

Young field: a slow start, then steep.

Your return on the next hour is the slope:

Factor F · Frontier
Conceptual illustration

In generative ML the frontier is months away, not a PhD away

Months, not a PhD.

Elsewhere: years of prerequisites before the first new result.

Why here: a young, mostly empirical field, open papers and tools, and questions small enough for a student budget.

Factor C · Compounding

The same math shows up in every model

Factor S · Scarcity
Conceptual illustration

Few people hold the whole stack, and that is the scarce part

Research: ask a sharp question, design the test.

Engineering: build systems that work.

Infrastructure: make it run on real machines.

Theory: know why it works.

Many people have one. Few have all four.

Factor P · Proof

Leave with proof other people can read

“I took a course” is a claim. A paper, a demo and a repo are evidence.

Return on effort

Low return hides inside ML too

ProjectSkills it trainsAnswer known?StakesYou can show
Tutorial cloneoneyes, in the tutorialnonea repo like everyone's
API wrapperone (glue code)yesnonesomeone else's model
Kaggle fine-tuneone (tuning)metric fixed in advancea ranka leaderboard number
This projectresearch, engineering, infra, theoryno: it's openclient, funding, a paperpaper, demo, code
Return on effort

This project scores on every factor

DDemand
3Foundations, not wrappers
6Real infrastructure
FFrontier
1A real open question
2Make-or-break moments
CCompounding
4How research actually works
5Serious engineering
SScarcity
7You own the whole stack
8Funded, with a client
PProof
9Something to show
effort
10It's a game: the hours go faster
The payoff

What you'll be able to say in a year

“I trained a diffusion model from scratch.”
“I designed a controlled generalization study.”
“I ran it on cloud GPUs.”
“I co-authored a paper.”CUCAI + workshop target
“I shipped a demo you can play.”

The cost: it's hard, and it's a year. None of this is guaranteed. All of it is reachable.

PART B

The project in one picture

A tiny game engine, a painter, and a diffusion model that learns to take over the painting.

The question

Chatforce's dream is huge; our question is precise

Our research question

When a diffusion renderer is trained on only some combinations of scene factors, where does its ability to render the unseen ones break, and why?

64×64 frames3 factors, 4–5 values eachan exact answer key

Inspired by GameNGen: a diffusion model that simulates Doom.

GameNGen: Valevski et al., “Diffusion Models Are Real-Time Game Engines”, arXiv:2408.14837 (2024), ICLR 2025.

The system

Every frame flows through four stages

The engine runs the game: map, player, enemies. It never draws the final picture.

One ray per screen column: how far to the wall, what it hit, where on the wall.

Channels: pixel-aligned maps of what is where. Structure, no style.

The skinner paints a themed frame from the channels. Same input, same pixels, every time.

Our diffusion model learns to do the skinner's job, starting from noise.

The skinner stays on as the answer key for every test.

The system · themes

A theme is data, not code

Theme record · 1 of 4

Dungeon

patternbricks, flagstonesfinishmattefogbrown, dense
Theme record · 2 of 4

Frozen

patternice tiles, blocksfinishglossyfogpale blue, thick
Theme record · 3 of 4

Neon

patternpanels, gridfinishemissive trimfogblack, dense
Theme record · 4 of 4

Overgrown

patternmossy stonefinishmattefoggreen mist
Same channels, four themes

Adding a theme = adding a row

So we can hand-write themes or sample as many as we want.

The question, as a picture

The research question is a grid with holes in it

Trained on: most of the pairings.

Can the model render the dark cells?

Every theme and every lighting was seen. Only the pairing is new.

And we know the exact right answer for every one.

Then sweep the coverage and find where it breaks. Deck 6 goes deep.

The team

Five streams, and each of you owns one

Stream 1

Engine & rendering

The ray-caster, map generator and channel outputs.

Stream 2

Skinning & data

The theme space, the skinner, dataset generation.

Stream 3

Model & training

The U-Net, conditioning and sampler. Train it.

Stream 4

Evaluation & analysis

Holdout splits, metrics, the curves in the paper.

Stream 5

Infra & MLOps

AWS jobs, storage, experiment tracking.

Today

Today: from a single neuron to our system

Tomorrow

Tomorrow, you teach it back

Explaining it is how you find out what you don't understand yet.

Elapsed
0:00:00
Steps on this slide
Next slide