Why this is worth a year of your time, and what we are going to build.
Same input. Very different output.
The difference is where you point it.
How much the skill is valued right now.
Hours until you can do something new.
Reused in the next thing you learn.
How few people have the combination.
Do you leave with evidence.
API wrapper: a quick win, then flat.
Mature field: the easy wins are taken.
Young field: a slow start, then steep.
Your return on the next hour is the slope:
Elsewhere: years of prerequisites before the first new result.
Why here: a young, mostly empirical field, open papers and tools, and questions small enough for a student budget.
Research: ask a sharp question, design the test.
Engineering: build systems that work.
Infrastructure: make it run on real machines.
Theory: know why it works.
Many people have one. Few have all four.
“I took a course” is a claim. A paper, a demo and a repo are evidence.
| Project | Skills it trains | Answer known? | Stakes | You can show |
|---|---|---|---|---|
| Tutorial clone | one | yes, in the tutorial | none | a repo like everyone's |
| API wrapper | one (glue code) | yes | none | someone else's model |
| Kaggle fine-tune | one (tuning) | metric fixed in advance | a rank | a leaderboard number |
| This project | research, engineering, infra, theory | no: it's open | client, funding, a paper | paper, demo, code |
The cost: it's hard, and it's a year. None of this is guaranteed. All of it is reachable.
A tiny game engine, a painter, and a diffusion model that learns to take over the painting.
When a diffusion renderer is trained on only some combinations of scene factors, where does its ability to render the unseen ones break, and why?
Inspired by GameNGen: a diffusion model that simulates Doom.
GameNGen: Valevski et al., “Diffusion Models Are Real-Time Game Engines”, arXiv:2408.14837 (2024), ICLR 2025.
The engine runs the game: map, player, enemies. It never draws the final picture.
One ray per screen column: how far to the wall, what it hit, where on the wall.
Channels: pixel-aligned maps of what is where. Structure, no style.
The skinner paints a themed frame from the channels. Same input, same pixels, every time.
Our diffusion model learns to do the skinner's job, starting from noise.
The skinner stays on as the answer key for every test.
So we can hand-write themes or sample as many as we want.
Trained on: most of the pairings.
Every theme and every lighting was seen. Only the pairing is new.
And we know the exact right answer for every one.
Then sweep the coverage and find where it breaks. Deck 6 goes deep.
The ray-caster, map generator and channel outputs.
The theme space, the skinner, dataset generation.
The U-Net, conditioning and sampler. Train it.
Holdout splits, metrics, the curves in the paper.
AWS jobs, storage, experiment tracking.
Explaining it is how you find out what you don't understand yet.