QMIND Research · Deck 7 of 7
PyTorch, the compute stack and MLOps
How our code trains, where it runs, what it costs, and how every piece of the system fits together.
A PyTorch, deep
B Training code → serving
C Renting compute
D The system
E Next steps
~30 sec
Last deck of the day. We went from one neuron to diffusion; this one is about the machines and the code that make our project real.
Five parts: what PyTorch really is; getting a trained model out and running; renting GPUs on AWS without wasting the credits; the whole system design; and what we do at the next meeting and tomorrow.
Intuition: research ideas are cheap; a pipeline that reliably turns an idea into a measured curve is what makes a team fast.
If asked: "Do I need all of this to contribute?" No: each stream owns one slice. But everyone should know the shape of the whole system so the interfaces between slices line up.
PART A
PyTorch, from the inside
It looks like NumPy. Underneath, it records derivatives and launches GPU kernels.
~15 sec
Matt's framing: PyTorch is "just" an auto-differentiation library you call like NumPy. True, and the next eleven slides open up what that sentence actually means.
Intuition: once you can picture what a tensor is in memory and what autograd records, every PyTorch error message starts to make sense.
PyTorch · what it is
PyTorch is NumPy with three superpowers
import numpy as npW = np.random.randn(3 , 16 ) x = np.random.randn(64 , 3 ) h = np.maximum(x @ W, 0 )
import torchW = torch.randn(3 , 16 , device="cuda" , requires_grad=True ) x = torch.randn(64 , 3 , device="cuda" ) h = torch.relu(x @ W) h.sum().backward() # fills W.grad
Same n-dimensional arrays, same broadcasting, almost the same names.
1 · Speed
GPU execution The same op runs as a kernel across thousands of GPU cores.
2 · Calculus
Automatic differentiation Every op is recorded so .backward() can compute every gradient.
3 · Toolkit
A neural-net library nn, optim, data loading, compilation, multi-GPU.
~1 min
Start with NumPy: an n-dimensional array plus fast vectorised math. A tiny layer: inputs times weights, then ReLU.
The PyTorch version is almost line-for-line identical. That is deliberate: PyTorch copied NumPy's array model and broadcasting rules.
Superpower 1: device="cuda" . The tensor lives in GPU memory, and every op launches a GPU kernel instead of running on the CPU.
Superpower 2: requires_grad=True tells PyTorch to record every operation on W. Then backward() walks that record and fills W.grad . This is the heart of PyTorch, and the next three slides.
Superpower 3: everything around it: layers, optimizers, DataLoader, torch.compile, distributed training. That is what makes it a framework rather than a math library.
Intuition: NumPy computes values; PyTorch computes values and remembers how it computed them , so it can differentiate.
Analogy: NumPy is a calculator; PyTorch is a calculator with a tape recorder, plus a garage of GPUs.
If asked: "Why not JAX?" JAX is excellent and also does autodiff + accelerators, but it is functional and traces programs; PyTorch is eager and imperative, which makes debugging with print and pdb easier. Most research code, and every tutorial we'll use, is PyTorch.
PyTorch · tensors in memory
A tensor is a view: storage, shape, strides
Indexing is arithmetic on metadata, not a lookup.
x.t() just swaps the strides. Same storage, zero copies.
.contiguous() pays for a real copy in the new order.
Broadcasting is stride 0: one row reused, never copied.
~1.5 min
Memory is one flat line of numbers. A tensor's data lives in a storage : here six floats, back to back.
The tensor itself is tiny metadata on top: shape (2, 3) and strides (3, 1). Stride = how many storage slots you jump to move one step along that dimension. Next row: jump 3. Next column: jump 1.
So element (1, 2) is at 1×3 + 2×1 = 5. No search, just a multiply-add.
Transpose: shape becomes (3, 2), strides become (1, 3). Watch: the storage did not move at all; only the mapping changed. Element 5 is now at (2, 1): 2×1 + 1×3 = 5. That's why .t() , .view() , slicing and .expand() are free.
Some kernels need rows to be adjacent in memory. .contiguous() copies into a fresh storage in the new order, strides (2, 1). This is where the classic error "view size is not compatible... use .reshape()" comes from: reshape copies if it must.
Broadcasting: a (3,) vector expanded to (2, 3) gets stride 0 on the first dimension: every row points at the same three numbers. Adding it to a matrix costs no extra memory.
Intuition: a tensor = pointer + shape + strides + dtype + device. Most "reshaping" is just editing the metadata.
Analogy: a spreadsheet view over one long column of data: you can re-sort the view without touching the data.
If asked: "When does a view bite you?" In-place ops on a view modify the original, and autograd will complain if you overwrite something it saved for backward.
PyTorch · autograd
Autograd records the forward pass, then runs it backwards
a = torch.tensor(2.0 , requires_grad=True ) b = torch.tensor(-3.0 , requires_grad=True ) c = torch.tensor(10.0 , requires_grad=True ) e = a * b d = e + c f = torch.tensor(-2.0 , requires_grad=True ) L = d * f L.backward() # a.grad == 6.0
Deck 2's chain rule: upstream gradient × local gradient, node by node.
~2.5 min
A tiny example (it's the one from Karpathy's micrograd). Scalars, so we can follow every number.
Leaves: a, b, c have requires_grad=True. They're the "parameters" of this toy.
e = a·b. PyTorch computes −6 and attaches a node, MulBackward0 , that remembers a and b, because it will need them to differentiate.
d = e + c = 4, node AddBackward0. The graph is built while the code runs : that's a "dynamic graph". An if-statement or loop in Python just builds a different graph.
L = d·f = −8. The forward pass is done, and we have a full record of how L was made.
backward(): start at L with gradient 1. A multiply node sends each input the upstream gradient times the other input: d gets 1·f = −2, f gets 1·d = 4.
An add node passes the gradient through unchanged: e and c both get −2.
Last multiply: a gets −2·b = 6, b gets −2·a = −4. Every leaf now has .grad filled.
This is exactly the chain rule from deck 2's backprop calculus. The only new thing is that a machine does the bookkeeping: each node knows its local derivative, and gradients are multiplied along the path.
Intuition: forward pass = write the recipe down; backward pass = read it in reverse, multiplying local slopes.
If asked: "What if a variable is used twice?" Its gradients from each use are added (multivariable chain rule). That's also why .grad accumulates across backward calls, and why we call zero_grad().
PyTorch · why reverse mode
Reverse mode: one backward pass gets every gradient
Forward mode
One sweep per input weight
1
passes
Reverse mode (backprop)
One sweep from the loss back
1
pass, about 2× the cost of a forward pass
One scalar loss, weights: we want all of at once. Reverse mode's cost doesn't grow with .
~1 min
Picture a network as a funnel: many weights on the left, one loss on the right.
Forward-mode AD pushes a derivative forward from one input at a time: one sweep tells you how depends on , another for … For a 30-million-parameter U-Net that is 30 million sweeps per step. Hopeless.
Reverse mode starts from the single output and pushes sensitivities backwards; one sweep reaches every weight. That's backprop.
Rule of thumb: a backward pass costs about twice a forward pass, so a training step ≈ 3× the forward FLOPs. We'll use that "3×" when we size GPUs in Part C.
Intuition: the cheap direction is from the side with fewer things. Many inputs, one output → go backwards.
If asked: "What's the catch?" Memory. Reverse mode must keep the forward activations until backward uses them. That's why activations, not weights, dominate training memory, and why tricks like activation checkpointing exist.
PyTorch · optimizers
The optimizer turns gradients into steps: SGD to Adam
opt = torch.optim.AdamW( model.parameters(), lr=2e-4 )
Adam stores and for every weight: two extra copies of the model.
~1.5 min
A long, narrow valley: steep across, shallow along. Real loss surfaces look like this in many directions at once.
Plain SGD: step against the gradient. The learning rate has to be small enough not to explode in the steep direction, so it bounces across the valley and crawls along it.
Adam keeps two running averages per weight: m, the average gradient (momentum, smooths the zigzag), and v, the average squared gradient (how big this weight's gradients usually are). Dividing by √v gives every weight its own step size. The hats are bias corrections for the first few steps, when m and v start at zero.
In code it's one line. AdamW = Adam with weight decay applied correctly (decoupled). It's the default for diffusion U-Nets and transformers. Note the memory: m and v are each as big as the model.
Intuition: SGD uses one step size for everyone; Adam gives each weight a step size scaled to how noisy and large its gradients are.
Analogy: SGD is a ball with no mass on an icy slope; Adam is a heavy ball with per-direction brakes.
If asked: "Is Adam always better?" No. Well-tuned SGD with momentum still wins on some vision tasks. But Adam is far less sensitive to the learning rate, which matters when you're running many sweeps.
PyTorch · the training loop
Every training run is the same five lines, repeated
model = UNet(cfg).cuda() opt = torch.optim.AdamW(model.parameters(), lr=2e-4 ) for x0, cond in loader: x0, cond = x0.cuda(), cond.cuda() t, noise = sample_t(x0), torch.randn_like(x0) opt.zero_grad() pred = model(add_noise(x0, noise, t), t, cond) loss = F.mse_loss(pred, noise) loss.backward() opt.step()
zero_grad · Clear last step's gradients. They accumulate by default.
forward · Run the U-Net on the noisy frame; autograd records the graph.
loss · How far the predicted noise is from the true noise (deck 5).
backward · Fill .grad for every one of the weights.
step · AdamW nudges every weight, in place.
Repeat around times. The curve is the only thing you watch.
~1.5 min
This is our actual diffusion training loop, minus logging. Top: build the model on the GPU and the optimizer. Then loop over batches. sample_t and add_noise are the forward-noising process from deck 5.
zero_grad: because gradients add up across backward calls (remember: used-twice variables add), we must clear them each step. Forgetting this is a classic bug: the effective learning rate grows every step.
Forward: noisy frame + timestep + conditioning channels in, predicted noise out. Autograd records the whole U-Net as it runs.
Loss: mean squared error between predicted and true noise: the "predict the noise" objective.
Backward: one reverse sweep fills .grad for all ~30M weights.
Step: the optimizer reads .grad and updates the weights in place.
That's it, ~100k times. Everything else in this deck (DataLoader, bf16, compile, checkpoints, AWS) exists to make these five lines run faster, safer, and somewhere else.
Intuition: training is a for-loop around "guess, measure the error, blame the weights, nudge them".
If asked: "Where do model.train() and model.eval() fit?" They switch layers like dropout and batch-norm between training and inference behaviour. Call eval() plus torch.no_grad() when sampling or evaluating.
PyTorch · data loading
The DataLoader keeps the GPU fed
loader = DataLoader(ds, batch_size=64 , shuffle=True , num_workers=8 , pin_memory=True )
A Dataset returns one sample by index; the DataLoader batches, shuffles and parallelises. Low GPU utilisation? Look here first.
~1 min
The GPU can only train as fast as batches arrive. Each sample must be read from disk, decoded, maybe augmented, and stacked into a batch: CPU work.
With num_workers=0, the main process does all of that, then hands the batch to the GPU, then waits. Look at the timeline: the GPU sits idle while the CPU loads. Very common in first projects.
With worker processes, several CPUs prepare batches in parallel and keep a small prefetch queue full. The GPU never waits. pin_memory puts batches in page-locked RAM so the copy to the GPU is faster and can overlap with compute.
The code. Dataset implements __len__ and __getitem__; DataLoader does the rest. Rule of thumb: start num_workers near the number of CPU cores, and watch nvidia-smi: if GPU utilisation is far below ~90%, the input pipeline is your bottleneck.
Intuition: the GPU is a fast chef; the DataLoader is the prep cooks. One prep cook can't keep up.
If asked: "Our frames are tiny, does this matter?" Yes, more so: at 64×64 the GPU finishes each batch quickly, so loading overhead is a bigger fraction. That's one reason we'll store data in shards and cache them on local disk.
PyTorch · devices and CUDA
Python only queues GPU work; .item() makes it wait
Kernel
One function run by thousands of GPU threads at once: a matmul, an add, a ReLU.
Asynchronous
Python enqueues kernels and moves on. The GPU drains the queue in order.
Syncs
.item(), .cpu(), print(t) wait for the GPU. Log every N steps, not every step.
~1 min
Two timelines: the CPU running your Python, and the GPU's work queue (a "CUDA stream").
Each PyTorch op on a CUDA tensor becomes a kernel launch : Python tells the driver "run matmul on these pointers" and returns in microseconds. The GPU executes kernels in order, each one using thousands of threads.
Because launches return immediately, Python runs ahead of the GPU. That's good: the queue stays full. It also means a Python timer around a GPU op measures almost nothing unless you call torch.cuda.synchronize().
Anything that needs the actual value on the CPU forces a sync: .item(), .cpu(), .numpy(), printing a tensor, or Python control flow on a tensor value. The CPU stalls until the GPU catches up, and the queue drains. Calling loss.item() every step can noticeably slow a small model.
Intuition: the CPU is a manager writing tickets; the GPU is the factory floor. Asking "what's the number right now?" makes the manager stand there waiting.
If asked: "Why so many tiny kernels?" Eager PyTorch launches one kernel per op. For small models, launch overhead and memory trips dominate, which is exactly what torch.compile attacks, two slides from now.
PyTorch · mixed precision
bf16 keeps float32's range with half the bits
sign exponent → range mantissa → precision
with torch.autocast("cuda" , dtype=torch.bfloat16): loss = F.mse_loss(model(x_t, t, cond), noise) loss.backward()
Matmuls run in bf16 on tensor cores; weights and sensitive sums stay fp32. Roughly 2× faster, half the activation memory.
~1 min
A float is three fields: sign, exponent (how big or small a number can be) and mantissa (how many significant digits).
float32: 8 exponent bits, 23 mantissa bits. The safe default, but 4 bytes per number and slower tensor-core math.
float16 halves the size but shrinks the exponent to 5 bits: anything above 65,504 overflows to infinity and tiny gradients underflow to zero. Training in fp16 needs "loss scaling" to dodge that.
bfloat16 ("brain float", from Google) keeps all 8 exponent bits and drops mantissa instead. It is literally the top 16 bits of a float32. Same range, so no loss scaling; less precision, which neural nets tolerate well.
autocast runs matmuls and convolutions in bf16 while keeping weights, the optimizer and reductions like the loss in fp32. All the GPUs we'd rent (A10G, L4, L40S, H100) support bf16 tensor cores.
Intuition: for gradients, being in the right ballpark (range) matters more than the fifth digit (precision).
If asked: "Will it hurt quality?" Usually not measurably for training. If we ever see NaNs, the first test is to rerun a few hundred steps in fp32 and compare; that's in our failure-points checklist.
PyTorch · torch.compile
torch.compile fuses small kernels into big ones
Eager: every op is its own kernel, and every intermediate makes a round trip to GPU memory.
Compiled: one fused kernel reads once, keeps intermediates in registers, writes once.
model = torch.compile(model)
TorchDynamo captures your Python as a graph → TorchInductor generates fused Triton kernels. First call is slow (it compiles); new input shapes can recompile.
~1 min
Take y = gelu(x·w + b), applied element-wise to a big activation tensor.
In eager mode that's three kernels. Each reads its input from GPU memory (HBM) and writes its output back. Six memory trips. Element-wise ops do almost no math per byte, so they are memory-bandwidth-bound: the same "memory-bound" idea as LLM decode in deck 4.
Compiled: the three ops are fused into one kernel. Read x once, do all three ops in registers, write y once. Two trips instead of six.
torch.compile (since PyTorch 2.0) does this automatically. Dynamo hooks Python bytecode to capture a graph of tensor ops; Inductor fuses and generates GPU code in Triton. Typical speedups range from none to well over 1.5×; it depends on the model, so we measure.
Intuition: the GPU is rarely slow at arithmetic; it's slow at fetching. Fusion means fetching less.
Analogy: three separate trips to the fridge vs. taking everything out at once.
If asked: "Any gotchas?" Graph breaks (Python Dynamo can't trace, e.g. printing, data-dependent control flow) split the graph and lose some benefit. Debug in eager, compile once things work. PyTorch 2.14 (released Sept 2, 2026; patch 2.14.1 on Sept 30) is the current stable release.
PyTorch · multi-GPU
DDP: a model copy per GPU, gradients averaged each step
model = DDP(model.cuda(rank), device_ids=[rank]) # launch: torchrun --nproc_per_node=4 train.py
Deck 4's data parallelism. Our U-Net fits on one GPU; DDP is for when we want results faster.
~1 min
DistributedDataParallel: one process per GPU, each holding a full copy of the model.
Each step, the global batch is split: each GPU gets its own slice of frames.
Each GPU runs forward and backward on its slice, so their gradients differ.
All-reduce: the GPUs exchange gradients (often in a ring) so every GPU ends up with the average. DDP overlaps this with the backward pass, bucket by bucket, so communication hides behind compute.
Every GPU applies the identical averaged update, so the replicas stay identical without ever sending weights. This is data parallelism from deck 4; tensor and pipeline parallelism are for models too big for one GPU, which ours is not.
Intuition: four students each grade a quarter of the exams, then agree on one average correction to the answer key.
If asked: "Should we use 8 GPUs?" Only if wall-clock matters more than cost. For sweeps, eight separate single-GPU runs (different seeds or settings) usually beat one 8-GPU run: no communication, and we need many runs anyway.
PART B
From training code to a running model
Save it so you can resume, get it out of the training script, and serve it fast enough to play.
~15 sec
We can train. Now: how a trained model leaves the training script and ends up rendering frames for someone holding a keyboard.
Intuition: training produces a file; serving is everything needed to turn that file into answers on demand.
Serving · checkpoints
A checkpoint is everything needed to resume, not just weights
torch.save({ "model" : model.state_dict(), "ema" : ema.state_dict(), "opt" : opt.state_dict(), "step" : step, "sched" : sched.state_dict(), "rng" : {"torch" : torch.get_rng_state(), "cuda" : torch.cuda.get_rng_state_all()}, "config" : cfg, "git" : GIT_SHA, "data" : DATASET_ID, }, f"ckpt/step_{step:07d}.pt" ) # then upload to S3
Resume = load all of it, then continue at step + 1. Test it: a resumed run should track the uninterrupted one.
Sizes are estimates for a ~30M-parameter U-Net stored in fp32 (4 bytes per number).
~1 min
A checkpoint is a Python dict saved with torch.save. The mistake is saving only the weights.
model.state_dict(): a dict from parameter name to tensor. For ~30M parameters in fp32, about 120 MB.
Diffusion models usually keep an EMA (exponential moving average) copy of the weights, which samples better; another 120 MB. Adam's m and v are each model-sized too. Without optimizer state, a resumed run gets a "cold" Adam and the loss jumps.
The small stuff is what makes it reproducible: step counter and LR schedule; RNG states so the noise sequence continues; and the config, git commit and dataset ID, so we can always answer "what exactly produced this model?"
Resuming: load everything, set the step, continue. Write a test that kills a run and resumes it. Spot instances (Part C) make this non-optional.
Intuition: a checkpoint is a save-game, not a screenshot. It must contain the whole world state.
If asked: "Will a resumed run be bit-identical?" Close but often not exact: the DataLoader's shuffle position and some non-deterministic GPU kernels also matter. Aim for "statistically the same curve", and log enough to tell.
Serving · export
Getting a model out of Python: four paths, one default
Path What you ship Status, Oct 2026 For us
state_dict + Pythonweights + the model class the standard training, eval, demo v0
torch.exportan ahead-of-time graph of the model recommended successor to TorchScript if we need Python-free inference
ONNX a portable graph file mature; runs on ONNX Runtime, TensorRT optimised runtimes, other languages
TorchScript a scripted or traced module deprecated: "use torch.export instead" don't start here
Rule: stay in Python until a measured need pushes you out.
Sources: PyTorch docs, TorchScript pages (deprecation notice); PyTorch 2.14 release announcement, Sept 2026.
~1 min
Once trained, how does the model leave the training script? Four options you'll see online.
The default: save the state_dict, and at inference import the same model class and load it. Simple, flexible, and the model code is version-controlled anyway. This is what we use.
torch.export: traces the model ahead of time into a clean graph of operations (an ExportedProgram) that other tools and runtimes can consume without your Python code.
ONNX: an open, cross-framework graph format. Use it to run on ONNX Runtime or compile with NVIDIA TensorRT for maximum inference speed, or to call from C++/C#/JS.
TorchScript was the old way; PyTorch's docs now say it's deprecated in favour of torch.export. Lots of old tutorials still use it: ignore them.
For a research project, staying in Python is right until profiling says otherwise. If the live demo is too slow, ONNX + TensorRT or torch.compile are the first levers.
Intuition: exporting means freezing the model's computation into data, so something other than your Python script can run it.
If asked: "Why not pickle the whole model object?" torch.save(model) pickles code references; it breaks when you rename or move a file. Saving the state_dict is robust.
Serving · the inference server proposal
A small inference server: channels in, frame out
A micro-batcher waits a few milliseconds to group requests: the GPU does 8 frames almost as fast as 1.
Throughput climbs with batch size; latency climbs too. Pick for your use.
Demo: FastAPI + one model process. Heavier traffic: NVIDIA Triton. TorchServe is no longer maintained.
Curve is illustrative. TorchServe status: PyTorch serve docs ("no longer actively maintained").
~1 min
An inference server is a process that holds the model in GPU memory and answers requests. Ours: the game sends structural channels plus a theme, the server returns a rendered frame.
One request: arrives over HTTP or a WebSocket, goes through a queue, runs a few denoising steps on the GPU, the frame comes back.
Many requests: a micro-batcher collects whatever arrives within a few milliseconds and runs them as one batch. A small U-Net at 64×64 doesn't fill the GPU with one frame, so batching is nearly free throughput.
The trade-off: larger batches mean more frames per second overall but each request waits longer. Offline rendering (evaluation videos) wants throughput; a player wants latency.
Proposal: a FastAPI app with one model process for the demo; it's 100 lines. If we ever need production-grade dynamic batching and multiple models, NVIDIA Triton Inference Server. TorchServe's docs now say it's no longer actively maintained, so skip it.
Intuition: a GPU is a bus, not a taxi. Filling the seats is cheap; making riders wait for the bus is the cost.
If asked: "Why a server at all for a local demo?" It cleanly separates the game process from the model process, and the same server renders evaluation videos in the cloud.
Serving · where it runs
Playable means the model sits next to the player
Cut : DDIM with 2–4 steps, later distillation. Cut : run the demo on a local GPU . Use the cloud to train and to render eval videos.
Timings are illustrative estimates; we will measure U-Net step time on our hardware. 30 fps = 33 ms per frame.
~1 min
Practical question: where does the engine run and where does the model run? Interactivity is a latency budget.
At 30 frames per second, every frame has 33 ms, end to end.
Local: the engine produces channels in about a millisecond, the U-Net runs K denoising steps (estimate a few ms each at 64×64 on a decent GPU, to be measured), the frame is displayed. Fits.
Cloud: before any compute, the request has to travel to a data centre and back; that alone is typically tens of milliseconds and it jitters. Add queueing and the same compute, and we blow the budget, every frame.
Two levers. K, the number of sampling steps: DDPM uses hundreds, DDIM can work with a handful, and distillation can reach 1–4 steps (a stretch goal). And the network: the demo model runs on a local GPU, a gaming laptop or a desktop in the room. Cloud GPUs are for training and offline rendering.
Intuition: you can buy more compute; you can't buy a shorter distance to Virginia.
If asked: "How did GameNGen do it?" It simulated Doom with a diffusion model at about 20 fps on a single TPU, using only a few DDIM sampling steps. Same lesson: few steps, compute right next to the game loop.
PART C
Renting compute
What our workload actually needs, what AWS charges for it, and how not to waste $5,000.
~15 sec
"How do you even go and rent H100s?" Let's answer that concretely, and also ask whether we should.
Intuition: the right GPU is the cheapest one that keeps our experiments moving, not the most famous one.
Compute · sizing the job
Our workload is small: one 24–48 GB GPU per run
GFLOP per 64×64 frame, : about FLOPs per run.
At ~30% of peak bf16 speed: hours per run.
Every GPU costs about the same per run . Cheap GPUs let us run sweeps in parallel; H100s only buy wall-clock time.
All workload numbers are rough estimates, to be measured. Peak dense bf16: L4 121, L40S 362 (NVIDIA datasheets via flopper.io), H100 SXM 989 TFLOPS. Prices: on-demand us-east-1, Oct 2026.
~1.5 min
Before renting anything: how big is our job? Back-of-envelope, clearly marked as estimates.
Training FLOPs ≈ 3 × forward FLOPs × samples seen (forward + backward ≈ 3×, from the reverse-mode slide). A ~30M-parameter U-Net at 64×64: on the order of 50 GFLOP per frame forward (estimate). 200k steps at batch 64 is ~13M samples. Total: about FLOPs.
Small conv nets rarely hit peak speed; assume ~30% utilisation. Then: L4 ~15 hours, L40S ~5 hours, H100 ~2 hours per run.
Multiply by hourly price: roughly $12, $9 and $12 per run. Nearly the same! The H100 is 3× faster and ~3.7× pricier than the L40S. And a small 64×64 model probably can't saturate an H100, so its real number is likely worse.
Memory check: weights + Adam + EMA ≈ 0.5 GB; activations at batch 64 a few GB. A 24 GB card is plenty, 48 GB is luxurious. Conclusion: single L4/A10G/L40S instances, many in parallel for sweeps. H100s only if a deadline needs one big run fast.
Intuition: for a sweep of 50 small runs, ten cheap GPUs beat one fast GPU.
If asked: "Where does 50 GFLOP come from?" Scaling from published DDPM CIFAR-10 U-Nets (~35M params at 32×32, a few GFLOPs) by 4× for 64×64 pixels, rounded up. Off by 2× either way wouldn't change the conclusion. Week 1 task: measure it with a profiler.
Compute · instance families
The AWS GPU menu, priced (us-east-1, on-demand)
Instance GPU GPU memory vCPUs $/hour
g6.xlarge 1× L4 24 GB 4 0.805
g5.xlarge 1× A10G 24 GB 4 1.006
g6e.xlarge 1× L40S 48 GB 4 1.861
p5.4xlarge 1× H100 80 GB 16 6.88
p4d.24xlarge 8× A100 8 × 40 GB 96 21.96
p5.48xlarge 8× H100 8 × 80 GB 192 55.04
p5en.48xlarge 8× H200 8 × 141 GB 192 63.30
p6-b200.48xlarge 8× B200 ~1.4 TB total 192 113.93
Linux on-demand, us-east-1, as of Oct 2026 (Economize, DevZero, Holori, DoiT price trackers; M-Star AWS GPU overview, Jul 2026). AWS cut P4d and P5 prices by up to 33% and 45% in June 2025. Prices change; check before launching.
~1 min
AWS names: the letter is the family (G = graphics/inference-class GPUs, P = top training GPUs), the number is the generation, the suffix is the size.
G family, one GPU each: g6 = NVIDIA L4 (24 GB, low power), g5 = A10G (24 GB, older Ampere), g6e = L40S (48 GB, ~3× an L4's bf16 throughput). Under $2/hour. Bars are price, linear scale: these barely register.
p5.4xlarge: a single H100 (80 GB). AWS added this one-GPU size, so you no longer have to rent eight H100s at once. About $6.88/hour.
The 8-GPU boxes: A100s, H100s, H200s (P5en; P5e also exists), and Blackwell B200s at over $100/hour (newer still: P6-B300, and G7e with RTX PRO 6000 Blackwell GPUs since Jan 2026). These are for training big models with fast GPU-to-GPU links, not for us.
Our sweet spot: g6.xlarge for development and cheap long runs, g6e.xlarge when we want speed or bigger batches.
Intuition: the price bars are the whole lesson: the jump from G to P is two orders of magnitude.
If asked: "What's a vCPU here?" A hyperthread. It matters because quotas (next slide) are counted in vCPUs, and our DataLoader workers need CPU too; g6.xlarge's 4 vCPUs can be tight, g6.2xlarge has 8.
Compute · how renting works
Renting a GPU on AWS, step by step
01
Account & credits One project account. Redeem the $5k credits. Turn on budget alerts first.
02
Quota request GPU quotas count vCPUs and start at 0 . Ask for "Running On-Demand G and VT instances". Can take days.
03
Region us-east-1: low prices, most GPUs. Keep S3 and EC2 in the same region; transfer between them is free.
04
Image A Deep Learning AMI (driver, CUDA, PyTorch installed) or our own Docker image.
05
Launch Pick g6e.xlarge, attach an IAM role, open SSH only to your IP.
06
Connect SSH or SSM Session Manager, VS Code Remote; run inside tmux.
07
Storage EBS disk for code and environment. S3 for datasets and checkpoints.
08
Stop vs terminate Stop keeps the disk (still billed). Terminate deletes it. Never leave a GPU idle.
AWS EC2 docs, "Amazon EC2 instance type quotas" (Oct 2026): G/VT and P quotas, On-Demand and Spot, default to 0 vCPUs.
~1.5 min
One AWS account for the project; redeem Chatforce's credits on it. Before launching anything, set an AWS Budget with email alerts.
The surprise for everyone: new accounts can't launch any GPU. Quotas are in vCPUs, and "Running On-Demand G and VT instances" and "Running On-Demand P instances" both start at 0. Spot has separate quotas, also 0. Request through Service Quotas; e.g. 32 vCPUs of G = eight g6e.xlarge at once. A p5.48xlarge alone needs 192. Approval can take a while, so this is a week-one ticket.
Region: us-east-1 (N. Virginia) has the most GPU capacity and among the lowest prices. Put the S3 bucket there too: S3 to EC2 in the same region costs nothing; cross-region and internet egress cost money.
AMI = machine image. AWS's Deep Learning AMI ships NVIDIA drivers, CUDA and PyTorch. Later we'll run our own Docker image so everyone's environment is identical.
Launch: instance type, an IAM role (so the box can read S3 without keys on disk), and a security group that only allows SSH from your IP.
Connect with SSH or SSM Session Manager (no open port needed). Run training inside tmux so it survives your laptop closing.
EBS = the instance's network disk; S3 = durable object storage for anything that must outlive the instance.
Stop = powered off; you stop paying for the GPU but still pay for the EBS disk. Terminate = gone, disk included (by default). Forgetting either is the #1 way students burn credits.
If asked: "Can we just click around in the console?" For the first instance, yes. After that, a launch script in infra/ so every run is reproducible and tagged.
Compute · pricing models
On-demand, spot, or a capacity block
On-demand · $6.88/hr
Start and stop any time, if capacity exists.
Spot · from ~$2.7/hr
Spare capacity, price moves. Can be reclaimed with a 2-minute warning.
Capacity Block · ~$5.19/hr
Reserved, prepaid: 1–14 days or whole weeks up to 182, booked up to 8 weeks ahead.
Proposal: on-demand G for dev, spot G for sweeps with frequent checkpoints, a block only for a deadline.
p5.4xlarge (1× H100), us-east-1, Oct 2026: on-demand and spot (Holori), Capacity Block rate (AWS Capacity Blocks pricing page). Spot curve is illustrative. AWS EC2 docs: spot interruption notices, Capacity Block durations.
~1.5 min
Same hardware, three ways to pay. Prices here are for one H100 (p5.4xlarge) because that's where the differences are biggest.
On-demand: the list price, billed per second. You can be refused with "insufficient capacity" if the region is out of that GPU.
Spot: AWS's unused capacity at a discount; trackers show H100 spot as low as ~$2.7/hr, varying by availability zone and hour. The catch: AWS can reclaim it with a two-minute warning (via instance metadata or EventBridge). The run bar shows why checkpoints matter: you only lose work since the last save, then resume on a new instance.
Capacity Blocks: reserve GPUs for a fixed window, prepaid at a price set when you book. Durations: 1–14 days, or multiples of 7 up to 182; start up to 8 weeks out. Instances are shut down 30 minutes before the block ends. Good for "we need 8 H100s for the paper deadline".
Our plan (proposal): an on-demand G instance for interactive development; spot G instances for sweeps, checkpointing every ~10–15 minutes; a capacity block only if a deadline demands H100s.
Intuition: spot is flying standby: cheap, and you'll occasionally get bumped.
If asked: "Do credits cover spot?" Credits generally apply to usage charges including spot; they don't cover upfront payments for Reserved Instances or Savings Plans. Capacity Blocks are also paid up front, and AWS's credit terms exclude "any upfront fee", so don't assume credits cover a block; confirm with AWS before booking one.
Compute · our budget
What $5,000 of credits buys
In runs (estimate)
At ~$9–12 per training run, that's a few hundred runs, minus dev time and mistakes.
Proposal: 60% sweeps · 20% development · 20% reserve.
Stretch it
Alliance (Compute Canada) clusters are free for researchers; undergrads join via a faculty sponsor. GPU clouds: H100 ~$3–4/hr.
Credits expire, stay with one account, and don't cover Marketplace or upfront reservations.
GPU-hours = $5,000 ÷ on-demand price (us-east-1, Oct 2026). Alliance sponsored-user roles: alliancecan.ca. GPU cloud rates: Lambda and RunPod pricing pages, Oct 2026. Credit terms: AWS Promotional Credit Terms (aws.amazon.com/awscredits).
~1 min
Chatforce gave us $5,000 in AWS credits. Let's turn dollars into GPU-hours.
On G instances: about 6,200 L4-hours, 5,000 A10G-hours or 2,700 L40S-hours. That is a lot of compute for a 64×64 model.
One H100 (p5.4xlarge): about 730 hours.
An 8×H100 box: about 91 hours, under four days. A B200 box: 44 hours. Same dollars, very different feel.
In runs: at our ~$9–12 per run estimate, a few hundred full runs. Realistically we lose a chunk to dev boxes, debugging and forgotten instances, so the proposal reserves 20%.
Free compute exists: the Digital Research Alliance of Canada (formerly Compute Canada) runs national clusters free for academic research; undergrads get accounts sponsored by a faculty PI. Worth asking a Queen's prof. GPU clouds like Lambda or RunPod list H100s around $3–4/hr. And read the credit fine print: AWS credits expire, are tied to the account, and don't pay Marketplace or reservation upfront fees.
Intuition: our constraint is not money for compute; it's discipline in not wasting it.
If asked: "Can we pay for Colab/Kaggle instead?" Fine for the MNIST/CIFAR warm-ups, but sessions time out and storage is ephemeral, so not for the real pipeline.
Compute · cost hygiene
The expensive mistake is the instance you forgot
p5.48xlarge left on, Friday 6 pm → Monday 6 am
$0
60 h × $55.04 · share of our $5,000
g6e.xlarge forgotten for a month
$0
720 h × $1.861
Auto-shutdown: an idle check that stops the box; jobs that end themselves.
Budgets: AWS Budgets alerts at 25 / 50 / 80% of credits.
Tags: every resource gets owner and project.
Clean up: terminate, then delete orphaned disks and snapshots.
No keys in Git. Use IAM roles and SSO. Bots scan public repos for AWS keys.
~1 min
The most common way student teams lose their cloud budget isn't a big experiment. It's something left running.
Someone launches an 8×H100 box to "try something" Friday evening and forgets it until Monday: about $3,300, two-thirds of our credits, for nothing.
Small mistakes add up too: one forgotten L40S instance for a month is over $1,300.
Guardrails: an idle-shutdown script on every dev box (e.g. if GPU utilisation is ~0 for 30 minutes, shut down), and training jobs that terminate their own instance when done. AWS Budgets emails us at thresholds.
Tag every instance, volume and bucket with owner and project so the weekly cost review can see who owns what. Stopped instances still bill for their disks; orphaned volumes and snapshots quietly accumulate.
Never commit access keys. Automated scanners look for AWS keys in public repos, and leaked keys get used to mine crypto on your account. Use IAM roles on instances and SSO for humans; if a key leaks, deactivate it immediately.
Intuition: the meter runs whether or not the GPU is doing anything.
If asked: "Who checks?" The Infra stream owns a weekly 5-minute cost review in Cost Explorer, grouped by the owner tag.
PART D
The system
From a seed to a rendered frame: every stage, every handoff, and where it breaks.
~15 sec
Now we zoom out and design the whole thing as a system, because five streams are going to build five pieces that must fit together.
Intuition: good system design is mostly about clean handoffs: what each stage promises the next one.
System design · architecture proposal
The whole system, assembled one stage at a time
frames & shards
checkpoints
metrics
rendered frames
~1.5 min
One diagram that every stream should be able to draw from memory. We'll build it left to right.
① Emulation: the game engine runs headless (no window) on CPUs, many copies at once, each from a seed. Output: structural channels per frame.
② Dataset generation: the skinner paints ground-truth frames for each theme and lighting; a writer packs samples into shard files plus a manifest, uploaded to S3 under a versioned dataset ID.
③ Training: a GPU job reads a config from Git and pulls shards from S3.
④ Tracking: metrics stream to the experiment tracker (W&B or MLflow), checkpoints go back to S3. Nothing important lives only on an instance.
⑤ Evaluation: separate jobs load checkpoints and render the held-out cells of the factor grid, computing PSNR, LPIPS and the factor classifiers (deck 6), logged next to the training run.
⑥ Demo: the best checkpoint goes to the local inference server; a player sends inputs, the engine makes channels, the model renders the frame. Packets are now moving everywhere: that's the system running.
Intuition: S3 is the hub. Every stage reads from or writes to it, so stages can run on different machines, at different times, owned by different people.
If asked: "Why not one big script?" Because the stages scale differently (CPU vs GPU), fail differently, and are owned by different streams. Storage in the middle decouples them.
System design · ① emulation
Run the engine headless: many copies, all seeded
# engine/headless.py (sketch) def render_episode (seed, n_frames): rng = np.random.default_rng(seed) world = make_map(rng) for _ in range(n_frames): world.step(policy(rng)) yield world.channels()
No window, no GPU: NumPy ray casting on CPU.
One process per core, on laptops or cheap CPU instances.
Same seed → bit-identical frames.
~1 min
"Emulating" the game for data means running it without a screen or a human. The engine is just a function: seed in, a stream of channel frames out. The frames on the right come from a real little ray-caster running in this slide, showing the depth channel.
Headless: the ray-caster never opens a window; it fills arrays. A scripted or random policy plays the game. No GPU needed; at 64×64, ray casting is cheap.
Parallelism is trivial: every episode is independent, so run one process per CPU core (Python multiprocessing), on our laptops or on a cheap CPU instance. This is "embarrassingly parallel".
Determinism: all randomness flows from the seed, so re-running seed 1003 gives bit-identical frames. We check this by hashing the arrays, live here. It makes datasets reproducible, bugs replayable, and lets evaluation regenerate any frame exactly.
Intuition: the engine is a pure function. Seeds are the dataset's DNA.
If asked: "What breaks determinism?" Unseeded global RNGs (random, np.random), iterating over Python sets, time-based logic, floating-point differences across machines. Pass an explicit rng everywhere, and make the CI test catch it.
System design · ② dataset generation
A dataset is shards plus a manifest, with a version
{ "id" : "ds-v003" , "channel_spec" : "0.1.0" , "git" : "a41c9e2" , "seeds" : [0 , 4000 ], "grid" : {"theme" :5 , "material" :5 , "light" :4 }, "shards" : [ {"file" : "shard-00000.tar" , "n" : 1000 , "sha256" : "9c1f…" }, … ] }
A million samples ≈ 53 GB ≈ per month in S3. Storage is cheap; GPU time is not.
S3 Standard us-east-1: $0.023/GB-month, PUT $0.005 per 1,000 requests; S3 → EC2 same region free (nubbo.app 2026 S3 pricing summary). Sample size is our estimate: RGB + f16 depth, uv, light + u8 ids.
~1 min
Generation workers take engine channels, run the skinner for each theme and lighting setting, and produce samples: ground-truth frame + channels + factor labels.
We pack samples into shards: tar or npz files of ~1,000 samples (~50 MB) each. Millions of tiny files are slow to list and copy, and cost real money in S3 requests: a million PUTs is $5 and painful to sync. Shards stream and cache well.
A manifest makes the dataset self-describing: an ID, the channel spec version, the generator's git commit and seed range, the factor grid, and every shard with a checksum. The ID is what a training config refers to.
Upload to S3 under the ID. Treat datasets as immutable: never overwrite ds-v003; fix a bug by generating ds-v004. Then any result can be traced to exact data.
Size: 64×64 pixels × ~13 bytes (RGB 3, depth 2, surface 1, material 1, uv 4, light 2) ≈ 53 KB per sample. A million samples is ~53 GB, about $1.20/month to store. Compare one GPU-hour. Storage is never our bottleneck.
Intuition: a dataset is a versioned artifact, like a software release, not a folder someone keeps editing.
If asked: "Where do the train/test splits live?" Not in the shards: a split file lists which factor cells are held out, versioned with the experiment. The same dataset serves many holdout schemes (deck 6).
System design · ③ ④ training and tracking
A training job: config in; checkpoints and metrics out
# configs/exp/coverage_50.yaml data: {dataset: ds-v003, holdout: random_15} model: {width: 128 , theme_enc: param_vector} train: {steps: 200000 , batch: 64 , lr: 2e-4 , precision: bf16, ckpt_every_min: 15 } coverage: 0.5 seed: 1
python train.py -m exp=coverage_50 seed=1 ,2 ,3
Where jobs run: scripted EC2 launches first. SageMaker training jobs add managed spot and auto-shutdown for a premium (ml.g5.xlarge ≈ $1.41/hr vs $1.01 on EC2).
Hydra-style config (proposal). SageMaker vs EC2 prices as reported by wring.co, 2026; verify on the SageMaker pricing page.
~1 min
A run is fully described by a config file in Git: which dataset ID and holdout split, model width and theme encoding, training settings, seed. If it's not in the config, it doesn't change between runs.
With Hydra (or plain YAML + argparse), one command launches a sweep: here, three seeds of the 50%-coverage experiment. Sweeps from deck 6 are just lists of configs.
Getting data in: stream shards straight from S3 (starts immediately, but every epoch hits the network), or sync them once to the instance's local NVMe disk and read at disk speed. At ~50 GB, sync-then-cache is simpler and faster for us.
Getting results out: every ~15 minutes a checkpoint goes to S3; every few steps, loss and learning rate go to the tracker, plus sample images. The tracker also stores the config, git commit and dataset ID, so every curve is traceable.
Where jobs run: first, a launch script that starts an EC2 instance, runs the job and terminates it. SageMaker training jobs do this for you (spin up, run, upload, shut down, with managed spot and resume), but ml.* instances cost more than the same EC2 box. Worth it once we're running many sweeps.
Intuition: a training job is a pure function too: (code commit, config, dataset ID) → (checkpoints, metrics).
If asked: "W&B or MLflow?" W&B: hosted, great dashboards, free academic tier. MLflow: open source, self-hosted. Either is fine; pick one in week one and make every run log to it.
System design · feature engineering
For us, feature engineering means channel design
target (skinner)
depth
surface
material
uv
light
Material ID: one-hot planes (exact, fixed) or a learned embedding (compact).
Depth: log-normalise so near detail isn't crushed and far walls don't saturate.
UV wraps from 1 back to 0; sin/cos keeps neighbours close.
Each choice is a dial in the encoding sweep (hypothesis 2). The channel spec pins it down.
~1.5 min
Classic ML has "feature engineering": choosing and transforming inputs. Our version: deciding exactly how each structural channel is encoded before it's concatenated with the noisy image.
These are real channels from a toy ray-caster running in the slide: the skinner's target, depth, surface type, material ID, texture coordinates, light.
Material ID is a category, not a number: material 4 isn't "twice" material 2. One-hot gives each material its own 0/1 plane: exact, no learned parameters, but M channels. An embedding (a learned vector per ID, applied per pixel) is compact and can learn similarities. That choice is an experiment.
Raw depth spans a huge range, so linear scaling crushes nearby detail. Log scaling to [0, 1] spreads it out. Networks train best on inputs of similar, bounded scale.
UV wraps: at a texture seam it jumps from 0.99 to 0.00, a fake edge. Encoding u as (sin 2πu, cos 2πu) makes 0.99 and 0.01 close again: the same trick as positional encodings in transformers (deck 3).
The input channel count follows from these choices. This ties directly to the encoding sweep: pixel-aligned channels vs a global ID vs a parameter vector. Whatever we pick is written down in the channel spec.
Intuition: the model can only use what we hand it, in the form we hand it. Good encodings make the right answer easy to compute.
If asked: "Why not let the network figure it out?" It can, given enough data and capacity, but we're testing generalization to unseen combinations; a bad encoding (e.g. material as one number) can create a failure we'd mistake for a finding.
System design · repo architecture proposal
One repo, folders that mirror the streams
qmind-render/ ├── engine/ ray-caster, maps, headless ├── skinner/ themes → ground truth ├── data/ generation, shards, manifests ├── model/ U-Net, conditioning, samplers ├── train/ loop, checkpoints, resume ├── eval/ splits, metrics, classifiers ├── infra/ AWS scripts, Docker, budgets ├── configs/ Hydra YAML ├── contracts/ channel_spec, theme_schema ├── notebooks/ exploration, never imported └── tests/ determinism, shapes, overfit
# contracts/channel_spec.yaml version: 0.1.0 size: [64 , 64 ] depth: {dtype: f16, norm: log, max: 32 } surface: [ceiling, wall, floor, enemy] material: {dtype: u8, count: 5 } uv: {dtype: f16, wrap: true } light: {dtype: f16, range: [0 , 1 ]}
Contracts are the handoffs between streams. Change one → bump the version → regenerate data. Every manifest records the version it used.
~1 min
One GitHub monorepo for the project. One repo means one commit hash identifies the exact state of every stage, which is what we log with each dataset and run.
The data-producing side: engine (Engine & rendering stream), skinner and data (Skinning & data pipeline stream).
The learning side: model and train (Model & training stream), eval (Evaluation & analysis stream).
The glue: infra (Infrastructure & MLOps), configs, notebooks for exploration only (if code in a notebook matters, it moves into a package), and tests.
contracts/ holds the two interfaces that cross stream boundaries: the channel spec (what the engine promises the model and skinner) and the theme schema (what a theme record contains). Here's a v0 sketch of the channel spec.
Treat them like an API: semantic versioning, reviewed changes, and every dataset manifest records which version produced it. If the engine team swaps u and v, the version bump forces everyone to notice instead of silently training on garbage.
Intuition: folders are ownership; contracts are promises between owners.
If asked: "Monorepo vs separate repos?" Separate repos make cross-cutting changes (spec v0.2) a multi-repo dance. At our size, one repo with clear folders and CODEOWNERS is simpler.
System design · continuous integration
CI runs three cheap tests that catch expensive bugs
def test_engine_is_deterministic (): a = frames(render_episode(seed=7 , n_frames=32 )) b = frames(render_episode(seed=7 , n_frames=32 )) assert sha256(a) == sha256(b) def test_channels_match_spec (): x = next(render_episode(seed=1 , n_frames=1 )) validate(x, load_spec("contracts/channel_spec.yaml" )) def test_tiny_model_overfits_one_batch (): loss = train(cfg.tiny, data=one_batch(), steps=300 ) assert loss < 1e-2
~1 min
Continuous integration: on every push or pull request, GitHub Actions runs our tests on a CPU runner. Three tests buy most of the value.
Determinism: render the same seed twice and compare hashes. Catches an unseeded RNG the moment it's introduced, not three weeks later when a dataset can't be reproduced.
Contract: generated channels have the shapes, dtypes and ranges the spec promises. Catches spec drift between the engine and model streams.
Overfit one batch: a tiny U-Net must drive the loss near zero on a single batch within a few hundred steps. If it can't memorise one batch, something is wired wrong: the loss, the conditioning, the data/label alignment. It runs on CPU in about a minute (estimate).
When a change breaks a test, the PR is blocked. Here someone's engine change introduced nondeterminism; CI says no before it poisons a dataset.
Intuition: the overfit test is the single most useful sanity check in deep learning: a model that can't memorise can't learn.
If asked: "Should CI train real models?" No: CI must be fast and cheap. Real training is triggered by us, on AWS. CI just guarantees the code we launch is not obviously broken.
System design · failure points
Where this pipeline breaks, and the guard for each
Data
1 Nondeterminismseed everything, CI test
2 Channels misaligned with framestore them in one sample
3 Train/test leakagesplit by factor cell, assert
Training
4 Silent NaNsassert finite loss, log grad norm
5 GPU idle, loader slowworkers, local cache
6 Out of memorysmaller batch, bf16
Infrastructure
7 Spot loss, no checkpointsave every 15 min
8 Forgotten instancesauto-shutdown, budgets
9 Untracked configslog config + git + data ID
10 "Works on my machine"pinned Docker image
Evaluation
11 Eval bugs that flatter the modelsanity-check the metric on known answers
~1.5 min
The pipeline again, compressed. Now: where it actually breaks. Each failure gets a guard we build in from day one.
Data. (1) Nondeterminism: the dataset can't be regenerated, so bugs can't be reproduced. (2) Misalignment: a frame paired with the previous frame's channels, or a flipped axis; the model learns garbage that "almost works". Write channels and RGB in the same sample; visually overlay them in a notebook. (3) Leakage: a held-out combination sneaks into training and inflates generalization. Split by factor cell and assert the sets are disjoint.
Training. (4) NaNs that silently turn the model to noise: check the loss is finite every step, log gradient norms. (5) The GPU idles because data loading is slow: watch utilisation. (6) Out of memory: reduce the batch or use gradient accumulation; bf16 halves activations.
Infrastructure. (7) A spot interruption with no recent checkpoint loses hours. (8) Forgotten instances burn credits. (9) A great result whose config nobody recorded can't be reproduced. (10) "Works on my machine": pin package versions and run in one Docker image everywhere.
Evaluation. (11) The most dangerous kind, because it feels like success: evaluating on training cells, PSNR on the wrong pixel range, comparing to the wrong frame. Sanity-check every metric: ground truth vs itself should be perfect; a blurred or wrong-theme frame should score badly.
Intuition: bugs that crash are cheap; bugs that produce plausible numbers are expensive.
If asked: "What's the one habit that catches most of these?" Look at the data and the outputs with your own eyes, early and often: sample grids in the tracker every few thousand steps.
Next steps · recap
The whole day in one line
~1 min
Let's walk the whole day in one line.
A neuron: weighted sum plus a nonlinearity. Stack them and you get a function with thousands of knobs.
Backprop: the chain rule tells every knob which way to turn. Today we saw it's exactly what autograd automates.
Transformers: tokens, the residual stream, attention letting every token look at the others.
The frontier: the same idea at data-centre scale: KV caches, parallelism, RL, test-time compute.
Diffusion: learn to remove noise, and you can generate images, and render our game.
Our data: a deterministic engine and skinner, a factor grid with held-out combinations, so generalization is graded exactly.
Our stack: PyTorch, AWS, and a pipeline that turns seeds into curves. You now know every layer from the neuron to the cloud bill.
Intuition: it's one idea all the way down: a differentiable function, a loss, and gradients, scaled up and pointed at a question.
Next steps · next meeting proposed
Next meeting: from understanding to tickets
Everyone: sign up for a stream · set up the environment (repo, Python, PyTorch, tests pass) · agree channel spec v0 together
Engine
Engine skeleton Headless ray-caster emitting channels per spec v0, seeded, with the determinism test.
Skinning & data
Skinner v0 Two themes, brick and tile patterns; shard writer + manifest.
Model
Warm-up diffusion Tiny DDPM on MNIST, then CIFAR-10. Then the 64×64 U-Net.
Evaluation
Splits + metrics Holdout split code; PSNR/LPIPS harness, sanity-checked.
Infra & MLOps
AWS + repo Account, budgets, GPU quota request , repo + CI, tracker.
Do the quota request this week: it's the one task that waits on someone else.
~1 min
The goal of today was that at the next meeting we stop learning and start building. Here's the proposed first sprint.
Everyone: choose a stream, get the environment working (clone, install, tests pass on your machine), and together write channel spec v0, since it's the contract most streams depend on.
Engine: the headless skeleton that emits spec-v0 channels from a seed. Skinning & data: skinner v0 with two themes plus the shard writer.
Model: warm up with a tiny diffusion model on MNIST then CIFAR-10, the fastest way to make deck 5 concrete, then the real U-Net. Evaluation: split code and a metric harness that we've proven correct on known answers.
Infra: AWS account, budget alerts, quota request, repo skeleton with CI, tracker workspace.
The quota request has external latency, so it goes first, this week.
Intuition: first tickets are chosen to make the contracts real early, so streams can work in parallel.
If asked: "What if I want two streams?" Pick a primary; pairing across streams on the contracts is encouraged.
Next steps · tomorrow's mini-presentations
Tomorrow: teach one idea in five minutes
Format
5 minutes, 3 slides Teach it to a first-year One picture, one equation Why it matters for our project
Backprop & the chain ruledeck 2
Gradient descent & learning ratedeck 2
Attention: queries, keys, valuesdeck 3
The residual streamdeck 3
KV cache & quadratic costdeck 4
Data, tensor, pipeline parallelismdeck 4
RLHF vs GRPOdeck 4
Test-time computedeck 4
The forward diffusion processdeck 5
The U-Netdeck 5
Held-out compositionsdeck 6
Autograd: what backward() doesdeck 7
Claim a topic before you leave. One person per topic.
~1.5 min
Tomorrow everyone presents. Teaching something is the fastest way to find out whether you understand it.
Format: five minutes, three slides. Pitch it at a first-year. One picture and one equation; end with why it matters for our project.
From deck 2 and 3: backprop, gradient descent, attention, the residual stream.
From deck 4: the KV cache and why attention is quadratic, the kinds of parallelism, RLHF vs GRPO, test-time compute.
From decks 5–7: forward diffusion, the U-Net, held-out compositions, and autograd from today.
Claim one before leaving so there are no duplicates. If the team is bigger than twelve, other good ones: bf16, renting GPUs on AWS, the failure points.
Intuition: you don't really know it until you can explain it simply.
If asked: "Can I use these slides?" Yes, borrow figures, but make your own explanation; that's the point.
QMIND Research · end of the day
Next time, we build it.
From one neuron to the machines that train our model. Tomorrow you teach it; next meeting we render the first frame.