Measured GPU-memory scaling — native / seasonal compute sizing¶
Generated by scripts/measure_memory_scaling.py. Peak memory of one DINN → integrate → loss → backward pass (eager autograd, pre-checkpointing) vs problem size. Memory of a fixed graph is value-independent, so synthetic inputs reproduce a real fit's footprint at matched shapes.
-
Device: NVIDIA GeForce RTX 5090 Laptop GPU (23.9 GB total)
-
Box: 15-tracer 2-layer, DT=0.25 d, DINN (per-cell 1x1 conv).
-
Fit
peak = a + b·(cells·steps): a = 4.5 MB, b = 356 B per cell·step, R² = 0.99999 (n=20).
Measurements¶
| cells | steps | train peak (GB) | forward-only (GB) | AD overhead x |
|---|---|---|---|---|
| 2,000 | 50 | 0.03 | 0.00 | 38.0 |
| 2,000 | 100 | 0.07 | 0.00 | 75.5 |
| 2,000 | 200 | 0.14 | 0.00 | 150.5 |
| 2,000 | 400 | 0.27 | 0.00 | 300.4 |
| 4,000 | 50 | 0.07 | 0.00 | 38.1 |
| 4,000 | 100 | 0.14 | 0.00 | 75.7 |
| 4,000 | 200 | 0.27 | 0.00 | 150.8 |
| 4,000 | 400 | 0.54 | 0.00 | 301.1 |
| 8,000 | 50 | 0.14 | 0.00 | 38.0 |
| 8,000 | 100 | 0.27 | 0.00 | 75.5 |
| 8,000 | 200 | 0.54 | 0.00 | 150.3 |
| 8,000 | 400 | 1.07 | 0.00 | 300.1 |
| 16,000 | 50 | 0.27 | 0.01 | 38.0 |
| 16,000 | 100 | 0.53 | 0.01 | 75.3 |
| 16,000 | 200 | 1.07 | 0.01 | 150.1 |
| 16,000 | 400 | 2.13 | 0.01 | 299.6 |
| 32,000 | 50 | 0.54 | 0.01 | 38.0 |
| 32,000 | 100 | 1.07 | 0.01 | 75.4 |
| 32,000 | 200 | 2.13 | 0.01 | 150.1 |
| 32,000 | 400 | 4.25 | 0.01 | 299.7 |
Extrapolation to the proposal grids¶
Seasonal assumes ~2000 steps (one annual cycle at DT=0.25 d ≈ 1460 + spin-up); time-mean ~200. Pre-checkpointing peak = the upper bound to size hardware against.
| target | cells | steps | predicted peak (GB) |
|---|---|---|---|
| LLC90 global · time-mean | 105,300 | 200 | 7.0 |
| LLC90 global · seasonal | 105,300 | 2000 | 69.9 |
| LLC270 global · time-mean | 947,700 | 200 | 62.9 |
| LLC270 global · seasonal | 947,700 | 2000 | 629.2 |
| 3-AOI native · seasonal (~20k cells) | 20,000 | 2000 | 13.3 |