Track 1 v2.2 — Phase 2: 5-PFT box-model extension¶
⚠ SUPERSEDED FRAMING (2026-06-28). Point-in-time record; data stands, framing corrected by STATUS.md. The project is a surrogate-to-model identifiability study over 4 observable params; the growth pair is unobservable by construction. R_PICPOC is recoverable with a real calcite anchor (the '5/6 ceiling / 6/6 wall / needs the Darwin port + native resolution' framing is refuted). The per-cell predictor is load-bearing for the target trio (ablation 7/10 vs 0/10, PR #158).
Status as of 2026-05-11: partial — best result is nb23 at 3/6 calibration-grade vs Carroll's published Green's-functions optima. v2.2.1 (per-PFT K_Fe, nb25) executing as of this writing.
The goal¶
Recover Carroll et al. 2020 / 2022 calibrated parameter values — the published Green's-functions optima encoded in CARROLL_VALUES — at calibration-grade quality (≤ 40% off Carroll) for all 6 Carroll-6 parameters. v2.0 hit calibration-grade on the iron pair only (alpfe, scav_rat); the other 4 (Smallgrow, Biggrow, diatomgraz, R_PICPOC) drifted because the 2-PFT (Ps + Pl) box averaged across 5 different Darwin phytoplankton functional types with very different rates.
The intervention¶
5-PFT box extension. Replace the lumped Ps (small phyto) + Pl (large phyto) with five distinct functional types matching Darwin 3 v05:
| State index | PFT | Empirical Eq Pacific abundance (surface) | Maps to |
|---|---|---|---|
| 1 | diatoms | 29.6% of Chl | Chl1, diatomgraz |
| 2 | other large eukaryotes | 21.5% of Chl | Chl2, Biggrow |
| 3 | Synechococcus | 1.6% of Chl | Chl3 |
| 4 | Pro low-light | ~0% of Chl (deep-adapted) | Chl4 |
| 5 | Pro high-light | 47.2% of Chl | Chl5, Smallgrow |
State vector grows from 7 to 10 tracers. Loss expands from 7 to 11 targets (FeT + 5 separate Chl_i + POC + PIC + DIC + ALK + CO2_flux). Each Carroll-6 parameter now governs one specific PFT instead of an average — Smallgrow learns Pro-HL's growth rate specifically; Biggrow learns other-large-euks specifically; diatomgraz learns diatom grazing specifically; the iron pair + R_PICPOC stay global.
Results¶
nb23 — 5-PFT box, Darwin targets, shared K_FE (v2.2 minimum-viable)¶
Branch v2.2-5pft-box commit c177a46. DINN baseline trained 1500 epochs.
| Param | recovered | Carroll publ. | |Δ|/Carroll | Band |
|---|---|---|---|---|
| alpfe | 0.101 | 0.928 | 0.891 | Loose |
| scav_rat | 4.22e-7 | 6.03e-7 | 0.300 | ✓ Calibration-grade |
| Smallgrow | 1.483 | 0.661 | 1.244 | Drifted |
| Biggrow | 0.291 | 0.431 | 0.326 | ✓ Calibration-grade |
| diatomgraz | 0.596 | 0.830 | 0.282 | ✓ Calibration-grade |
| R_PICPOC | 0.011 | 0.042 | 0.738 | Loose |
3 of 6 at calibration-grade. Phase 2 hypothesis partially confirmed: Biggrow and diatomgraz (previously Drifted in v2.0) moved to Calibration-grade; R_PICPOC improved from Drifted (3.6 off) to Loose (0.7 off). But alpfe regressed from v2.0's Excellent (1% off) to Loose (89% off) — suspected cause: shared K_FE aliasing across PFTs with very different iron physiology.
nb24 — 5-PFT box + GLODAP DIC/ALK real-obs combo (Phase 2 + Phase 1 stack)¶
Branch v2.2-5pft-box commit 7a938f8. DINN baseline trained 1500 epochs.
| Param | recovered | Carroll publ. | |Δ|/Carroll | Band |
|---|---|---|---|---|
| alpfe | 0.137 | 0.928 | 0.852 | Loose |
| scav_rat | 3.49e-7 | 6.03e-7 | 0.420 | Loose |
| Smallgrow | 1.413 | 0.661 | 1.137 | Drifted |
| Biggrow | 1.094 | 0.431 | 1.536 | Drifted |
| diatomgraz | 0.760 | 0.830 | 0.084 | ✓ Excellent |
| R_PICPOC | 0.120 | 0.042 | 1.833 | Drifted |
1 of 6 at calibration-grade. The combo did not stack — it conflicted. Worse than nb23 on scav_rat, Biggrow, R_PICPOC (all degraded out of Cal-grade or further into Drifted). The one win was striking: diatomgraz jumped from Cal-grade (0.282 off) to Excellent (0.084 off) — DINNDeep on the same run hit 0.005 off, the best single-parameter recovery on this project. GLODAP ALK is a stronger CaCO₃ constraint than Darwin's internal ALK, lighting up the diatom-grazing-driven calcification signal.
Mechanism: GLODAP DIC + ALK have different spatial structure than Darwin's internal self-consistent fields. The 5-PFT optimizer can't satisfy both the Darwin-Chl signal and the GLODAP-carbonate signal at the same parameter set, so it lands in a compromise region. Phase 2 + Phase 1 combo is not the v2.2 deliverable.
nb25 — 5-PFT box + per-PFT K_Fe half-saturations (v2.2.1)¶
Branch v2.2-5pft-box commit 3a11b7f (executed). DINN baseline trained 1500 epochs in 2509 s.
| Param | recovered | Carroll publ. | |Δ|/Carroll | Band |
|---|---|---|---|---|
| alpfe | 0.138 | 0.928 | 0.851 | Loose |
| scav_rat | 4.44e-7 | 6.03e-7 | 0.263 | ✓ Calibration-grade |
| Smallgrow | 1.043 | 0.661 | 0.577 | Loose |
| Biggrow | 0.814 | 0.431 | 0.888 | Loose |
| diatomgraz | 0.882 | 0.830 | 0.062 | ✓ Excellent |
| R_PICPOC | 0.005 | 0.042 | 0.882 | Loose |
2 of 6 at calibration-grade. v2.2.1 hypothesis REJECTED. Per-PFT K_FE did not restore alpfe (0.891 → 0.851 off — essentially flat). The shared-K_FE-aliasing hypothesis is wrong. Mixed-direction signal: diatomgraz improved from Cal-grade (0.282 off) to Excellent (0.062 off — diatom Fe-uptake-driven mortality is genuinely K_FE-sensitive); but Biggrow regressed from Cal-grade to Loose (literature-plausible large-euk K_FE = 60 nM was probably too high vs the effective average baked into Carroll's Biggrow calibration).
Working hypothesis for the actual root cause: the 11-target loss is 7:1 weighted toward carbonate/Chl signals (5 separate Chl_i + POC + PIC + DIC + ALK vs 1 FeT target), starving the iron-pair capacity. v2.2.2 candidate fix is loss weighting (upweight FeT) rather than further per-PFT differentiation.
nb23 (3/6 calibration-grade) is the best v2.2 result. v2.2.1 is a failed-hypothesis-but-informative experiment.
What v2.2 is and isn't¶
Is: the structural fix v2.0 and Phase 1 both identified as needed — separating phytoplankton functional types so each Carroll-6 parameter governs one specific species' dynamics. nb23 shows the principle works: 3 of the 4 previously-drifted params moved closer to Carroll, two into Calibration-grade.
Isn't: a clean "all 6 hit Carroll" headline. The shared-K_FE simplification carried an aliasing cost that pulled alpfe off-optimum. v2.2.1 (nb25, in flight) tests whether per-PFT K_Fe closes that gap.
Path forward¶
| Hypothesis test | Status | Result |
|---|---|---|
| 5-PFT box (vs 2-PFT) | Phase 2, nb23 | ✓ 3/6 cal-grade, partial confirmation |
| Phase 2 + GLODAP DIC/ALK combo | nb24 | ✗ Combo conflicts; combo strategy rejected for Eq Pacific |
| Per-PFT K_Fe half-saturations | v2.2.1, nb25 | ✗ Rejected — alpfe regression persists; diatomgraz went to Excellent, Biggrow regressed |
| Loss weighting (upweight FeT) | v2.2.2 candidate | Strongest current candidate — attacks the 7:1 carbonate:iron loss imbalance |
| Per-PFT Q_Fe (iron quotas) | v2.2.3 candidate | Untested alternative |
| GEOTRACES direct iron observations | Phase 3 (cluster-gated) | Not yet started |
v2.2 ships when all 6 params land in Calibration-grade or better on a single notebook run. Current best is nb29 (PINN drift w=3.0) at 4/6 (Biggrow, diatomgraz, scav_rat, R_PICPOC all calibration-grade). alpfe + Smallgrow are confirmed structurally stuck across 20 experiments. The remaining unblockers are Carroll's initial-condition fields (Jon Q2) and/or GEOTRACES real iron data; neither is achievable through loss-function engineering.
nb29 — 5-PFT box + PINN drift constraint, w=3.0 (v2.4 — current best)¶
Branch v2.2-5pft-box commit 23f8061. DINN baseline trained 1500 epochs in 174 s with torch.compile.
| Param | recovered | Carroll publ. | |Δ|/Carroll | Band |
|---|---|---|---|---|
| alpfe | 0.105 | 0.928 | 0.888 | Loose (unchanged) |
| scav_rat | 3.95e-7 | 6.03e-7 | 0.345 | ✓ Calibration-grade |
| Smallgrow | 1.485 | 0.661 | 1.251 | Drifted (unchanged) |
| Biggrow | 0.567 | 0.431 | 0.314 | ✓ Calibration-grade |
| diatomgraz | 0.583 | 0.830 | 0.299 | ✓ Calibration-grade |
| R_PICPOC | 0.058 | 0.042 | 0.358 | ✓ Calibration-grade |
4 of 6 at calibration-grade. PINN drift constraint pulled R_PICPOC into cal-grade (0.738 nb23 → 0.358) while keeping Biggrow, diatomgraz, scav_rat. The physical coupling worked for carbonate-iron dynamics.
What 20 overnight experiments confirmed about alpfe¶
| Intervention category | Weights tested | alpfe range |
|---|---|---|
| z-scored 11-target baseline | nb23 + 5 seeds | 0.841 – 0.891 |
| FeT pattern upweighting (z-scored) | FET_W = 3.0 | 0.833 |
| Raw FeT MSE (magnitude) | w = 0.01, 0.05, 0.1, 0.3, 0.5, 1.0, 3.0 | 0.392 – 0.935 |
| GLODAP DIC + ALK hybrid | Phase 1 | 0.852 |
| Per-PFT K_FE | nb25 | 0.851 |
| PINN balance (strict source-sink) | w = 0.3, 1.0 | 0.882 – 0.887 |
| PINN drift (relative rate-of-change) | w = 0.05, 3.0 | 0.888 – 0.893 |
| Combined raw-FeT + PINN drift | combos | 0.817 – 0.862 |
alpfe sat at 0.80–0.94 off Carroll under every intervention. The single exception was raw_fet w=0.01, which moved alpfe to 0.392 but broke scav_rat to 2.556 off (the iron-pair tradeoff swapped roles).
Confirmed structural ceiling. Methodology improvements alone cannot close the alpfe gap. The remaining hypotheses (each requires data we don't have):
- IC compensation — Carroll's published alpfe was tuned for specific 3D initial fields. Jon Q2 in tonight's email asks for those fields.
- Steady-state assumption violation — 50-day box integration is too short for iron pool equilibration. Longer integration may help PINN succeed.
- GEOTRACES real iron — provides the absolute magnitude constraint our z-scored loss strips and our Darwin-internal FeT doesn't preserve.
Key files¶
src/darwindiff/carroll6_5pft.py— 10-tracer 5-PFT box modulesrc/darwindiff/glodap_loader.py— cherry-picked from PR #36 branchtests/test_carroll6_5pft.py— 9 tests covering smoke, autograd, per-PFT mapping, per-PFT K_Fe pathnotebooks/23_5pft_box_eqpac.ipynb— Phase 2 with Darwin-only targets, shared K_FEnotebooks/24_5pft_box_glodap_hybrid_eqpac.ipynb— Phase 2 + GLODAP combo (rejected)notebooks/25_5pft_box_perfe_eqpac.ipynb— v2.2.1, per-PFT K_FE (in flight)scripts/phase2_p4_p5_check.py— reproducible P4 (Eq Pac PFT abundance) + P5 (VRAM budget)scripts/build_nbXX.py— nbformat builders for each notebook
Memory + thread state¶
feedback_recovery_comparison_framing.md— compare against Carroll-published, not inter-notebook deltasfeedback_drop_dinndeep_phase2plus.md— v2.2.x trains DINN baseline only; DINNDeep saturates and adds no recovery info- D: thread
v2.1-phase2-5pft-box— active, full chronology + open questions