Measured GPU-memory scaling — native / seasonal compute sizing¶
Generated by scripts/measure_memory_scaling.py. Peak memory of one DINN → integrate → loss → backward pass (torch.compile'd per-step, pre-checkpointing) vs problem size. Memory of a fixed graph is value-independent, so synthetic inputs reproduce a real fit's footprint at matched shapes.
-
Device: NVIDIA H200 (139.8 GB total)
-
Box: 15-tracer 2-layer, DT=0.25 d, DINN (per-cell 1x1 conv).
-
Fit
peak = a + b·(cells·steps): a = 13.1 MB, b = 82.9 B per cell·step, R² = 0.99485 (n=20).
Measurements¶
| cells | steps | train peak (GB) | forward-only (GB) | AD overhead x |
|---|---|---|---|---|
| 2,000 | 50 | 0.07 | 0.06 | 1.1 |
| 2,000 | 100 | 0.02 | 0.00 | 30.7 |
| 2,000 | 200 | 0.03 | 0.00 | 59.4 |
| 2,000 | 400 | 0.06 | 0.00 | 116.9 |
| 4,000 | 50 | 0.08 | 0.06 | 1.3 |
| 4,000 | 100 | 0.03 | 0.00 | 30.7 |
| 4,000 | 200 | 0.06 | 0.00 | 59.5 |
| 4,000 | 400 | 0.13 | 0.00 | 117.1 |
| 8,000 | 50 | 0.04 | 0.00 | 16.3 |
| 8,000 | 100 | 0.07 | 0.00 | 30.7 |
| 8,000 | 200 | 0.13 | 0.00 | 59.5 |
| 8,000 | 400 | 0.25 | 0.00 | 117.1 |
| 16,000 | 50 | 0.07 | 0.00 | 16.3 |
| 16,000 | 100 | 0.13 | 0.00 | 30.7 |
| 16,000 | 200 | 0.26 | 0.00 | 59.5 |
| 16,000 | 400 | 0.50 | 0.00 | 117.1 |
| 32,000 | 50 | 0.14 | 0.01 | 16.3 |
| 32,000 | 100 | 0.26 | 0.01 | 30.8 |
| 32,000 | 200 | 0.51 | 0.01 | 59.6 |
| 32,000 | 400 | 1.01 | 0.01 | 117.2 |
Extrapolation to the proposal grids¶
Seasonal assumes ~2000 steps (one annual cycle at DT=0.25 d ≈ 1460 + spin-up); time-mean ~200. Pre-checkpointing peak = the upper bound to size hardware against.
| target | cells | steps | predicted peak (GB) |
|---|---|---|---|
| LLC90 global · time-mean | 105,300 | 200 | 1.6 |
| LLC90 global · seasonal | 105,300 | 2000 | 16.3 |
| LLC270 global · time-mean | 947,700 | 200 | 14.6 |
| LLC270 global · seasonal | 947,700 | 2000 | 146.3 |
| 3-AOI native · seasonal (~20k cells) | 20,000 | 2000 | 3.1 |