QMIND Research · Deck 7 of 7

PyTorch, the compute stack and MLOps

How our code trains, where it runs, what it costs, and how every piece of the system fits together.

APyTorch, deep BTraining code → serving CRenting compute DThe system ENext steps
PART A

PyTorch, from the inside

It looks like NumPy. Underneath, it records derivatives and launches GPU kernels.

PyTorch · what it is

PyTorch is NumPy with three superpowers

import numpy as npW = np.random.randn(3, 16)x = np.random.randn(64, 3)h = np.maximum(x @ W, 0)
import torchW = torch.randn(3, 16, device="cuda", requires_grad=True)x = torch.randn(64, 3, device="cuda")h = torch.relu(x @ W)h.sum().backward()   # fills W.grad

Same n-dimensional arrays, same broadcasting, almost the same names.

1 · Speed

GPU execution

The same op runs as a kernel across thousands of GPU cores.

2 · Calculus

Automatic differentiation

Every op is recorded so .backward() can compute every gradient.

3 · Toolkit

A neural-net library

nn, optim, data loading, compilation, multi-GPU.

PyTorch · tensors in memory

A tensor is a view: storage, shape, strides

Indexing is arithmetic on metadata, not a lookup.

x.t() just swaps the strides. Same storage, zero copies.

.contiguous() pays for a real copy in the new order.

Broadcasting is stride 0: one row reused, never copied.

PyTorch · autograd

Autograd records the forward pass, then runs it backwards

a = torch.tensor(2.0, requires_grad=True)b = torch.tensor(-3.0, requires_grad=True)c = torch.tensor(10.0, requires_grad=True)e = a * bd = e + cf = torch.tensor(-2.0, requires_grad=True)L = d * fL.backward()   # a.grad == 6.0

Deck 2's chain rule: upstream gradient × local gradient, node by node.

PyTorch · why reverse mode

Reverse mode: one backward pass gets every gradient

Forward mode

One sweep per input weight

1

passes

Reverse mode (backprop)

One sweep from the loss back

1

pass, about 2× the cost of a forward pass

One scalar loss, weights: we want all of at once. Reverse mode's cost doesn't grow with .

PyTorch · optimizers

The optimizer turns gradients into steps: SGD to Adam

SGD

ADAM

opt = torch.optim.AdamW(    model.parameters(), lr=2e-4)

Adam stores and for every weight: two extra copies of the model.

PyTorch · the training loop

Every training run is the same five lines, repeated

model = UNet(cfg).cuda()opt = torch.optim.AdamW(model.parameters(), lr=2e-4)for x0, cond in loader:    x0, cond = x0.cuda(), cond.cuda()    t, noise = sample_t(x0), torch.randn_like(x0)    opt.zero_grad()    pred = model(add_noise(x0, noise, t), t, cond)    loss = F.mse_loss(pred, noise)    loss.backward()    opt.step()

zero_grad · Clear last step's gradients. They accumulate by default.

forward · Run the U-Net on the noisy frame; autograd records the graph.

loss · How far the predicted noise is from the true noise (deck 5).

backward · Fill .grad for every one of the weights.

step · AdamW nudges every weight, in place.

Repeat around times. The curve is the only thing you watch.

PyTorch · data loading

The DataLoader keeps the GPU fed

loader = DataLoader(ds, batch_size=64,    shuffle=True, num_workers=8, pin_memory=True)

A Dataset returns one sample by index; the DataLoader batches, shuffles and parallelises. Low GPU utilisation? Look here first.

PyTorch · devices and CUDA

Python only queues GPU work; .item() makes it wait

Kernel

One function run by thousands of GPU threads at once: a matmul, an add, a ReLU.

Asynchronous

Python enqueues kernels and moves on. The GPU drains the queue in order.

Syncs

.item(), .cpu(), print(t) wait for the GPU. Log every N steps, not every step.

PyTorch · mixed precision

bf16 keeps float32's range with half the bits

signexponent → rangemantissa → precision
float32
max · ~7 digits
float16
max · ~3 digits
bfloat16
max · ~2–3 digits
with torch.autocast("cuda", dtype=torch.bfloat16):    loss = F.mse_loss(model(x_t, t, cond), noise)loss.backward()

Matmuls run in bf16 on tensor cores; weights and sensitive sums stay fp32. Roughly 2× faster, half the activation memory.

PyTorch · torch.compile

torch.compile fuses small kernels into big ones

Eager: every op is its own kernel, and every intermediate makes a round trip to GPU memory.

Compiled: one fused kernel reads once, keeps intermediates in registers, writes once.

model = torch.compile(model)

TorchDynamo captures your Python as a graph → TorchInductor generates fused Triton kernels. First call is slow (it compiles); new input shapes can recompile.

PyTorch · multi-GPU

DDP: a model copy per GPU, gradients averaged each step

model = DDP(model.cuda(rank), device_ids=[rank])# launch:  torchrun --nproc_per_node=4 train.py

Deck 4's data parallelism. Our U-Net fits on one GPU; DDP is for when we want results faster.

PART B

From training code to a running model

Save it so you can resume, get it out of the training script, and serve it fast enough to play.

Serving · checkpoints

A checkpoint is everything needed to resume, not just weights

torch.save({  "model": model.state_dict(),  "ema":   ema.state_dict(),  "opt":   opt.state_dict(),  "step": step, "sched": sched.state_dict(),  "rng": {"torch": torch.get_rng_state(),          "cuda": torch.cuda.get_rng_state_all()},  "config": cfg, "git": GIT_SHA, "data": DATASET_ID,}, f"ckpt/step_{step:07d}.pt")   # then upload to S3

Resume = load all of it, then continue at step + 1. Test it: a resumed run should track the uninterrupted one.

Sizes are estimates for a ~30M-parameter U-Net stored in fp32 (4 bytes per number).

Serving · export

Getting a model out of Python: four paths, one default

PathWhat you shipStatus, Oct 2026For us
state_dict + Pythonweights + the model classthe standardtraining, eval, demo v0
torch.exportan ahead-of-time graph of the modelrecommended successor to TorchScriptif we need Python-free inference
ONNXa portable graph filemature; runs on ONNX Runtime, TensorRToptimised runtimes, other languages
TorchScripta scripted or traced moduledeprecated: "use torch.export instead"don't start here

Rule: stay in Python until a measured need pushes you out.

Sources: PyTorch docs, TorchScript pages (deprecation notice); PyTorch 2.14 release announcement, Sept 2026.

Serving · the inference server proposal

A small inference server: channels in, frame out

A micro-batcher waits a few milliseconds to group requests: the GPU does 8 frames almost as fast as 1.

Throughput climbs with batch size; latency climbs too. Pick for your use.

Demo: FastAPI + one model process. Heavier traffic: NVIDIA Triton. TorchServe is no longer maintained.

Curve is illustrative. TorchServe status: PyTorch serve docs ("no longer actively maintained").

Serving · where it runs

Playable means the model sits next to the player

Cut : DDIM with 2–4 steps, later distillation. Cut : run the demo on a local GPU. Use the cloud to train and to render eval videos.

Timings are illustrative estimates; we will measure U-Net step time on our hardware. 30 fps = 33 ms per frame.

PART C

Renting compute

What our workload actually needs, what AWS charges for it, and how not to waste $5,000.

Compute · sizing the job

Our workload is small: one 24–48 GB GPU per run

GFLOP per 64×64 frame, : about FLOPs per run.

At ~30% of peak bf16 speed: hours per run.

Every GPU costs about the same per run. Cheap GPUs let us run sweeps in parallel; H100s only buy wall-clock time.

All workload numbers are rough estimates, to be measured. Peak dense bf16: L4 121, L40S 362 (NVIDIA datasheets via flopper.io), H100 SXM 989 TFLOPS. Prices: on-demand us-east-1, Oct 2026.

Compute · instance families

The AWS GPU menu, priced (us-east-1, on-demand)

InstanceGPUGPU memoryvCPUs$/hour
g6.xlarge1× L424 GB40.805
g5.xlarge1× A10G24 GB41.006
g6e.xlarge1× L40S48 GB41.861
p5.4xlarge1× H10080 GB166.88
p4d.24xlarge8× A1008 × 40 GB9621.96
p5.48xlarge8× H1008 × 80 GB19255.04
p5en.48xlarge8× H2008 × 141 GB19263.30
p6-b200.48xlarge8× B200~1.4 TB total192113.93

Linux on-demand, us-east-1, as of Oct 2026 (Economize, DevZero, Holori, DoiT price trackers; M-Star AWS GPU overview, Jul 2026). AWS cut P4d and P5 prices by up to 33% and 45% in June 2025. Prices change; check before launching.

Compute · how renting works

Renting a GPU on AWS, step by step

01

Account & credits

One project account. Redeem the $5k credits. Turn on budget alerts first.

02

Quota request

GPU quotas count vCPUs and start at 0. Ask for "Running On-Demand G and VT instances". Can take days.

03

Region

us-east-1: low prices, most GPUs. Keep S3 and EC2 in the same region; transfer between them is free.

04

Image

A Deep Learning AMI (driver, CUDA, PyTorch installed) or our own Docker image.

05

Launch

Pick g6e.xlarge, attach an IAM role, open SSH only to your IP.

06

Connect

SSH or SSM Session Manager, VS Code Remote; run inside tmux.

07

Storage

EBS disk for code and environment. S3 for datasets and checkpoints.

08

Stop vs terminate

Stop keeps the disk (still billed). Terminate deletes it. Never leave a GPU idle.

AWS EC2 docs, "Amazon EC2 instance type quotas" (Oct 2026): G/VT and P quotas, On-Demand and Spot, default to 0 vCPUs.

Compute · pricing models

On-demand, spot, or a capacity block

On-demand · $6.88/hr

Start and stop any time, if capacity exists.

Spot · from ~$2.7/hr

Spare capacity, price moves. Can be reclaimed with a 2-minute warning.

Capacity Block · ~$5.19/hr

Reserved, prepaid: 1–14 days or whole weeks up to 182, booked up to 8 weeks ahead.

Proposal: on-demand G for dev, spot G for sweeps with frequent checkpoints, a block only for a deadline.

p5.4xlarge (1× H100), us-east-1, Oct 2026: on-demand and spot (Holori), Capacity Block rate (AWS Capacity Blocks pricing page). Spot curve is illustrative. AWS EC2 docs: spot interruption notices, Capacity Block durations.

Compute · our budget

What $5,000 of credits buys

In runs (estimate)

At ~$9–12 per training run, that's a few hundred runs, minus dev time and mistakes.

Proposal: 60% sweeps · 20% development · 20% reserve.

Stretch it

Alliance (Compute Canada) clusters are free for researchers; undergrads join via a faculty sponsor. GPU clouds: H100 ~$3–4/hr.

Credits expire, stay with one account, and don't cover Marketplace or upfront reservations.

GPU-hours = $5,000 ÷ on-demand price (us-east-1, Oct 2026). Alliance sponsored-user roles: alliancecan.ca. GPU cloud rates: Lambda and RunPod pricing pages, Oct 2026. Credit terms: AWS Promotional Credit Terms (aws.amazon.com/awscredits).

Compute · cost hygiene

The expensive mistake is the instance you forgot

p5.48xlarge left on, Friday 6 pm → Monday 6 am

$0

60 h × $55.04 · share of our $5,000

g6e.xlarge forgotten for a month

$0

720 h × $1.861

  • Auto-shutdown: an idle check that stops the box; jobs that end themselves.
  • Budgets: AWS Budgets alerts at 25 / 50 / 80% of credits.
  • Tags: every resource gets owner and project.
  • Clean up: terminate, then delete orphaned disks and snapshots.
  • No keys in Git. Use IAM roles and SSO. Bots scan public repos for AWS keys.
PART D

The system

From a seed to a rendered frame: every stage, every handoff, and where it breaks.

System design · architecture proposal

The whole system, assembled one stage at a time

frames & shards checkpoints metrics rendered frames
System design · ① emulation

Run the engine headless: many copies, all seeded

# engine/headless.py  (sketch)def render_episode(seed, n_frames):    rng = np.random.default_rng(seed)    world = make_map(rng)    for _ in range(n_frames):        world.step(policy(rng))        yield world.channels()

No window, no GPU: NumPy ray casting on CPU.

One process per core, on laptops or cheap CPU instances.

Same seed → bit-identical frames.

System design · ② dataset generation

A dataset is shards plus a manifest, with a version

{ "id": "ds-v003",  "channel_spec": "0.1.0",  "git": "a41c9e2", "seeds": [0, 4000],  "grid": {"theme":5, "material":5, "light":4},  "shards": [    {"file": "shard-00000.tar", "n": 1000,     "sha256": "9c1f…"}, … ] }

A million samples ≈ 53 GB ≈ per month in S3. Storage is cheap; GPU time is not.

S3 Standard us-east-1: $0.023/GB-month, PUT $0.005 per 1,000 requests; S3 → EC2 same region free (nubbo.app 2026 S3 pricing summary). Sample size is our estimate: RGB + f16 depth, uv, light + u8 ids.

System design · ③ ④ training and tracking

A training job: config in; checkpoints and metrics out

# configs/exp/coverage_50.yamldata:  {dataset: ds-v003, holdout: random_15}model: {width: 128, theme_enc: param_vector}train: {steps: 200000, batch: 64, lr: 2e-4,        precision: bf16, ckpt_every_min: 15}coverage: 0.5seed: 1
python train.py -m exp=coverage_50 seed=1,2,3

Where jobs run: scripted EC2 launches first. SageMaker training jobs add managed spot and auto-shutdown for a premium (ml.g5.xlarge ≈ $1.41/hr vs $1.01 on EC2).

Hydra-style config (proposal). SageMaker vs EC2 prices as reported by wring.co, 2026; verify on the SageMaker pricing page.

System design · feature engineering

For us, feature engineering means channel design

target (skinner)
depth
surface
material
uv
light

Material ID: one-hot planes (exact, fixed) or a learned embedding (compact).

Depth: log-normalise so near detail isn't crushed and far walls don't saturate.

UV wraps from 1 back to 0; sin/cos keeps neighbours close.

Each choice is a dial in the encoding sweep (hypothesis 2). The channel spec pins it down.

System design · repo architecture proposal

One repo, folders that mirror the streams

qmind-render/├── engine/      ray-caster, maps, headless├── skinner/     themes → ground truth├── data/        generation, shards, manifests├── model/       U-Net, conditioning, samplers├── train/       loop, checkpoints, resume├── eval/        splits, metrics, classifiers├── infra/       AWS scripts, Docker, budgets├── configs/     Hydra YAML├── contracts/   channel_spec, theme_schema├── notebooks/   exploration, never imported└── tests/       determinism, shapes, overfit
# contracts/channel_spec.yamlversion: 0.1.0size: [64, 64]depth:    {dtype: f16, norm: log, max: 32}surface:  [ceiling, wall, floor, enemy]material: {dtype: u8, count: 5}uv:       {dtype: f16, wrap: true}light:    {dtype: f16, range: [0, 1]}

Contracts are the handoffs between streams. Change one → bump the version → regenerate data. Every manifest records the version it used.

System design · continuous integration

CI runs three cheap tests that catch expensive bugs

def test_engine_is_deterministic():    a = frames(render_episode(seed=7, n_frames=32))    b = frames(render_episode(seed=7, n_frames=32))    assert sha256(a) == sha256(b) def test_channels_match_spec():    x = next(render_episode(seed=1, n_frames=1))    validate(x, load_spec("contracts/channel_spec.yaml")) def test_tiny_model_overfits_one_batch():    loss = train(cfg.tiny, data=one_batch(), steps=300)    assert loss < 1e-2
System design · failure points

Where this pipeline breaks, and the guard for each

Data
1Nondeterminismseed everything, CI test
2Channels misaligned with framestore them in one sample
3Train/test leakagesplit by factor cell, assert
Training
4Silent NaNsassert finite loss, log grad norm
5GPU idle, loader slowworkers, local cache
6Out of memorysmaller batch, bf16
Infrastructure
7Spot loss, no checkpointsave every 15 min
8Forgotten instancesauto-shutdown, budgets
9Untracked configslog config + git + data ID
10"Works on my machine"pinned Docker image
Evaluation
11Eval bugs that flatter the modelsanity-check the metric on known answers
Next steps · recap

The whole day in one line

Next steps · next meeting proposed

Next meeting: from understanding to tickets

Everyone: sign up for a stream · set up the environment (repo, Python, PyTorch, tests pass) · agree channel spec v0 together

Engine

Engine skeleton

Headless ray-caster emitting channels per spec v0, seeded, with the determinism test.

Skinning & data

Skinner v0

Two themes, brick and tile patterns; shard writer + manifest.

Model

Warm-up diffusion

Tiny DDPM on MNIST, then CIFAR-10. Then the 64×64 U-Net.

Evaluation

Splits + metrics

Holdout split code; PSNR/LPIPS harness, sanity-checked.

Infra & MLOps

AWS + repo

Account, budgets, GPU quota request, repo + CI, tracker.

Do the quota request this week: it's the one task that waits on someone else.

Next steps · tomorrow's mini-presentations

Tomorrow: teach one idea in five minutes

Format

  • 5 minutes, 3 slides
  • Teach it to a first-year
  • One picture, one equation
  • Why it matters for our project
Backprop & the chain ruledeck 2
Gradient descent & learning ratedeck 2
Attention: queries, keys, valuesdeck 3
The residual streamdeck 3
KV cache & quadratic costdeck 4
Data, tensor, pipeline parallelismdeck 4
RLHF vs GRPOdeck 4
Test-time computedeck 4
The forward diffusion processdeck 5
The U-Netdeck 5
Held-out compositionsdeck 6
Autograd: what backward() doesdeck 7

Claim a topic before you leave. One person per topic.

QMIND Research · end of the day

Next time, we build it.

From one neuron to the machines that train our model. Tomorrow you teach it; next meeting we render the first frame.

Elapsed
0:00:00
Steps on this slide
Next slide