Skip to content

Measured GPU-memory scaling — native / seasonal compute sizing

Generated by scripts/measure_memory_scaling.py. Peak memory of one DINN → integrate → loss → backward pass (torch.compile'd per-step, pre-checkpointing) vs problem size. Memory of a fixed graph is value-independent, so synthetic inputs reproduce a real fit's footprint at matched shapes.

  • Device: NVIDIA H200 (139.8 GB total)

  • Box: 15-tracer 2-layer, DT=0.25 d, DINN (per-cell 1x1 conv).

  • Fit peak = a + b·(cells·steps): a = 13.1 MB, b = 82.9 B per cell·step, R² = 0.99485 (n=20).

Measurements

cells steps train peak (GB) forward-only (GB) AD overhead x
2,000 50 0.07 0.06 1.1
2,000 100 0.02 0.00 30.7
2,000 200 0.03 0.00 59.4
2,000 400 0.06 0.00 116.9
4,000 50 0.08 0.06 1.3
4,000 100 0.03 0.00 30.7
4,000 200 0.06 0.00 59.5
4,000 400 0.13 0.00 117.1
8,000 50 0.04 0.00 16.3
8,000 100 0.07 0.00 30.7
8,000 200 0.13 0.00 59.5
8,000 400 0.25 0.00 117.1
16,000 50 0.07 0.00 16.3
16,000 100 0.13 0.00 30.7
16,000 200 0.26 0.00 59.5
16,000 400 0.50 0.00 117.1
32,000 50 0.14 0.01 16.3
32,000 100 0.26 0.01 30.8
32,000 200 0.51 0.01 59.6
32,000 400 1.01 0.01 117.2

Extrapolation to the proposal grids

Seasonal assumes ~2000 steps (one annual cycle at DT=0.25 d ≈ 1460 + spin-up); time-mean ~200. Pre-checkpointing peak = the upper bound to size hardware against.

target cells steps predicted peak (GB)
LLC90 global · time-mean 105,300 200 1.6
LLC90 global · seasonal 105,300 2000 16.3
LLC270 global · time-mean 947,700 200 14.6
LLC270 global · seasonal 947,700 2000 146.3
3-AOI native · seasonal (~20k cells) 20,000 2000 3.1