Skip to content

Track 1 v2.2 — Phase 2: 5-PFT box-model extension

⚠ SUPERSEDED FRAMING (2026-06-28). Point-in-time record; data stands, framing corrected by STATUS.md. The project is a surrogate-to-model identifiability study over 4 observable params; the growth pair is unobservable by construction. R_PICPOC is recoverable with a real calcite anchor (the '5/6 ceiling / 6/6 wall / needs the Darwin port + native resolution' framing is refuted). The per-cell predictor is load-bearing for the target trio (ablation 7/10 vs 0/10, PR #158).

Status as of 2026-05-11: partial — best result is nb23 at 3/6 calibration-grade vs Carroll's published Green's-functions optima. v2.2.1 (per-PFT K_Fe, nb25) executing as of this writing.

The goal

Recover Carroll et al. 2020 / 2022 calibrated parameter values — the published Green's-functions optima encoded in CARROLL_VALUES — at calibration-grade quality (≤ 40% off Carroll) for all 6 Carroll-6 parameters. v2.0 hit calibration-grade on the iron pair only (alpfe, scav_rat); the other 4 (Smallgrow, Biggrow, diatomgraz, R_PICPOC) drifted because the 2-PFT (Ps + Pl) box averaged across 5 different Darwin phytoplankton functional types with very different rates.

The intervention

5-PFT box extension. Replace the lumped Ps (small phyto) + Pl (large phyto) with five distinct functional types matching Darwin 3 v05:

State index PFT Empirical Eq Pacific abundance (surface) Maps to
1 diatoms 29.6% of Chl Chl1, diatomgraz
2 other large eukaryotes 21.5% of Chl Chl2, Biggrow
3 Synechococcus 1.6% of Chl Chl3
4 Pro low-light ~0% of Chl (deep-adapted) Chl4
5 Pro high-light 47.2% of Chl Chl5, Smallgrow

State vector grows from 7 to 10 tracers. Loss expands from 7 to 11 targets (FeT + 5 separate Chl_i + POC + PIC + DIC + ALK + CO2_flux). Each Carroll-6 parameter now governs one specific PFT instead of an average — Smallgrow learns Pro-HL's growth rate specifically; Biggrow learns other-large-euks specifically; diatomgraz learns diatom grazing specifically; the iron pair + R_PICPOC stay global.

Results

nb23 — 5-PFT box, Darwin targets, shared K_FE (v2.2 minimum-viable)

Branch v2.2-5pft-box commit c177a46. DINN baseline trained 1500 epochs.

Param recovered Carroll publ. |Δ|/Carroll Band
alpfe 0.101 0.928 0.891 Loose
scav_rat 4.22e-7 6.03e-7 0.300 ✓ Calibration-grade
Smallgrow 1.483 0.661 1.244 Drifted
Biggrow 0.291 0.431 0.326 ✓ Calibration-grade
diatomgraz 0.596 0.830 0.282 ✓ Calibration-grade
R_PICPOC 0.011 0.042 0.738 Loose

3 of 6 at calibration-grade. Phase 2 hypothesis partially confirmed: Biggrow and diatomgraz (previously Drifted in v2.0) moved to Calibration-grade; R_PICPOC improved from Drifted (3.6 off) to Loose (0.7 off). But alpfe regressed from v2.0's Excellent (1% off) to Loose (89% off) — suspected cause: shared K_FE aliasing across PFTs with very different iron physiology.

nb24 — 5-PFT box + GLODAP DIC/ALK real-obs combo (Phase 2 + Phase 1 stack)

Branch v2.2-5pft-box commit 7a938f8. DINN baseline trained 1500 epochs.

Param recovered Carroll publ. |Δ|/Carroll Band
alpfe 0.137 0.928 0.852 Loose
scav_rat 3.49e-7 6.03e-7 0.420 Loose
Smallgrow 1.413 0.661 1.137 Drifted
Biggrow 1.094 0.431 1.536 Drifted
diatomgraz 0.760 0.830 0.084 ✓ Excellent
R_PICPOC 0.120 0.042 1.833 Drifted

1 of 6 at calibration-grade. The combo did not stack — it conflicted. Worse than nb23 on scav_rat, Biggrow, R_PICPOC (all degraded out of Cal-grade or further into Drifted). The one win was striking: diatomgraz jumped from Cal-grade (0.282 off) to Excellent (0.084 off) — DINNDeep on the same run hit 0.005 off, the best single-parameter recovery on this project. GLODAP ALK is a stronger CaCO₃ constraint than Darwin's internal ALK, lighting up the diatom-grazing-driven calcification signal.

Mechanism: GLODAP DIC + ALK have different spatial structure than Darwin's internal self-consistent fields. The 5-PFT optimizer can't satisfy both the Darwin-Chl signal and the GLODAP-carbonate signal at the same parameter set, so it lands in a compromise region. Phase 2 + Phase 1 combo is not the v2.2 deliverable.

nb25 — 5-PFT box + per-PFT K_Fe half-saturations (v2.2.1)

Branch v2.2-5pft-box commit 3a11b7f (executed). DINN baseline trained 1500 epochs in 2509 s.

Param recovered Carroll publ. |Δ|/Carroll Band
alpfe 0.138 0.928 0.851 Loose
scav_rat 4.44e-7 6.03e-7 0.263 ✓ Calibration-grade
Smallgrow 1.043 0.661 0.577 Loose
Biggrow 0.814 0.431 0.888 Loose
diatomgraz 0.882 0.830 0.062 Excellent
R_PICPOC 0.005 0.042 0.882 Loose

2 of 6 at calibration-grade. v2.2.1 hypothesis REJECTED. Per-PFT K_FE did not restore alpfe (0.891 → 0.851 off — essentially flat). The shared-K_FE-aliasing hypothesis is wrong. Mixed-direction signal: diatomgraz improved from Cal-grade (0.282 off) to Excellent (0.062 off — diatom Fe-uptake-driven mortality is genuinely K_FE-sensitive); but Biggrow regressed from Cal-grade to Loose (literature-plausible large-euk K_FE = 60 nM was probably too high vs the effective average baked into Carroll's Biggrow calibration).

Working hypothesis for the actual root cause: the 11-target loss is 7:1 weighted toward carbonate/Chl signals (5 separate Chl_i + POC + PIC + DIC + ALK vs 1 FeT target), starving the iron-pair capacity. v2.2.2 candidate fix is loss weighting (upweight FeT) rather than further per-PFT differentiation.

nb23 (3/6 calibration-grade) is the best v2.2 result. v2.2.1 is a failed-hypothesis-but-informative experiment.

What v2.2 is and isn't

Is: the structural fix v2.0 and Phase 1 both identified as needed — separating phytoplankton functional types so each Carroll-6 parameter governs one specific species' dynamics. nb23 shows the principle works: 3 of the 4 previously-drifted params moved closer to Carroll, two into Calibration-grade.

Isn't: a clean "all 6 hit Carroll" headline. The shared-K_FE simplification carried an aliasing cost that pulled alpfe off-optimum. v2.2.1 (nb25, in flight) tests whether per-PFT K_Fe closes that gap.

Path forward

Hypothesis test Status Result
5-PFT box (vs 2-PFT) Phase 2, nb23 ✓ 3/6 cal-grade, partial confirmation
Phase 2 + GLODAP DIC/ALK combo nb24 ✗ Combo conflicts; combo strategy rejected for Eq Pacific
Per-PFT K_Fe half-saturations v2.2.1, nb25 ✗ Rejected — alpfe regression persists; diatomgraz went to Excellent, Biggrow regressed
Loss weighting (upweight FeT) v2.2.2 candidate Strongest current candidate — attacks the 7:1 carbonate:iron loss imbalance
Per-PFT Q_Fe (iron quotas) v2.2.3 candidate Untested alternative
GEOTRACES direct iron observations Phase 3 (cluster-gated) Not yet started

v2.2 ships when all 6 params land in Calibration-grade or better on a single notebook run. Current best is nb29 (PINN drift w=3.0) at 4/6 (Biggrow, diatomgraz, scav_rat, R_PICPOC all calibration-grade). alpfe + Smallgrow are confirmed structurally stuck across 20 experiments. The remaining unblockers are Carroll's initial-condition fields (Jon Q2) and/or GEOTRACES real iron data; neither is achievable through loss-function engineering.

nb29 — 5-PFT box + PINN drift constraint, w=3.0 (v2.4 — current best)

Branch v2.2-5pft-box commit 23f8061. DINN baseline trained 1500 epochs in 174 s with torch.compile.

Param recovered Carroll publ. |Δ|/Carroll Band
alpfe 0.105 0.928 0.888 Loose (unchanged)
scav_rat 3.95e-7 6.03e-7 0.345 ✓ Calibration-grade
Smallgrow 1.485 0.661 1.251 Drifted (unchanged)
Biggrow 0.567 0.431 0.314 ✓ Calibration-grade
diatomgraz 0.583 0.830 0.299 ✓ Calibration-grade
R_PICPOC 0.058 0.042 0.358 ✓ Calibration-grade

4 of 6 at calibration-grade. PINN drift constraint pulled R_PICPOC into cal-grade (0.738 nb23 → 0.358) while keeping Biggrow, diatomgraz, scav_rat. The physical coupling worked for carbonate-iron dynamics.

What 20 overnight experiments confirmed about alpfe

Intervention category Weights tested alpfe range
z-scored 11-target baseline nb23 + 5 seeds 0.841 – 0.891
FeT pattern upweighting (z-scored) FET_W = 3.0 0.833
Raw FeT MSE (magnitude) w = 0.01, 0.05, 0.1, 0.3, 0.5, 1.0, 3.0 0.392 – 0.935
GLODAP DIC + ALK hybrid Phase 1 0.852
Per-PFT K_FE nb25 0.851
PINN balance (strict source-sink) w = 0.3, 1.0 0.882 – 0.887
PINN drift (relative rate-of-change) w = 0.05, 3.0 0.888 – 0.893
Combined raw-FeT + PINN drift combos 0.817 – 0.862

alpfe sat at 0.80–0.94 off Carroll under every intervention. The single exception was raw_fet w=0.01, which moved alpfe to 0.392 but broke scav_rat to 2.556 off (the iron-pair tradeoff swapped roles).

Confirmed structural ceiling. Methodology improvements alone cannot close the alpfe gap. The remaining hypotheses (each requires data we don't have):

  1. IC compensation — Carroll's published alpfe was tuned for specific 3D initial fields. Jon Q2 in tonight's email asks for those fields.
  2. Steady-state assumption violation — 50-day box integration is too short for iron pool equilibration. Longer integration may help PINN succeed.
  3. GEOTRACES real iron — provides the absolute magnitude constraint our z-scored loss strips and our Darwin-internal FeT doesn't preserve.

Key files

  • src/darwindiff/carroll6_5pft.py — 10-tracer 5-PFT box module
  • src/darwindiff/glodap_loader.py — cherry-picked from PR #36 branch
  • tests/test_carroll6_5pft.py — 9 tests covering smoke, autograd, per-PFT mapping, per-PFT K_Fe path
  • notebooks/23_5pft_box_eqpac.ipynb — Phase 2 with Darwin-only targets, shared K_FE
  • notebooks/24_5pft_box_glodap_hybrid_eqpac.ipynb — Phase 2 + GLODAP combo (rejected)
  • notebooks/25_5pft_box_perfe_eqpac.ipynb — v2.2.1, per-PFT K_FE (in flight)
  • scripts/phase2_p4_p5_check.py — reproducible P4 (Eq Pac PFT abundance) + P5 (VRAM budget)
  • scripts/build_nbXX.py — nbformat builders for each notebook

Memory + thread state

  • feedback_recovery_comparison_framing.md — compare against Carroll-published, not inter-notebook deltas
  • feedback_drop_dinndeep_phase2plus.md — v2.2.x trains DINN baseline only; DINNDeep saturates and adds no recovery info
  • D: thread v2.1-phase2-5pft-box — active, full chronology + open questions