Skip to content

I now have everything needed. One note: STATUS.md line 66 says DINN is "~454 weights" while ml-2 computes 406 (SST-only)/438 (default) — I'll flag the count needs verification rather than asserting 406. I have all facts confirmed. Writing the report.


DarwinDiff main.tex — Area-Chair Revision Plan

⚠ SUPERSEDED FRAMING (2026-06-27). Point-in-time paper-prep doc; data stands, framing corrected by STATUS.md. The project is a surrogate-to-model identifiability study over 4 OBSERVABLE params {alpfe, scav_rat, diatomgraz, R_PICPOC} (growth pair unobservable by construction), not a 6/6 chase / 'parameter conservation' ceiling. R_PICPOC is recoverable with a real calcite anchor (Daniels/MODIS); the Darwin port + native resolution were tested and did not help. The surrogate gap is dimensional (the 0-D box homogenizes spatial structure), so identifiability comes from real absolute anchors.

Prepared for Lucas (Ziming) Qi and Jonathan Lauderdale. Read against the canonical data (aggregated_v3.1.json: 856 seeds / 86 configs; Carroll optima: alpfe 0.928, scav_rat 6.025e-7, R_PICPOC 0.04245, diatomgraz 0.830) and v3.1_closeout.md. No numbers below are invented — every figure is either from the manuscript, the JSON, or the closeout. Counts are settled per your brief; I do not re-flag 856/86.

The paper is honest, internally consistent on the headline numbers, and the science is sound. The fixes below make it defensible under adversarial review without changing a single result. The single highest-leverage theme: the title and abstract's last sentence over-claim ("structural… intrinsic") while the body under-claims (robust ceiling is really 4/6, not 5/6). Fixing that asymmetry is worth more than all the minor edits combined.


1. Top revisions, prioritized

Deduped across lenses, ordered by impact. "MF" = must-fix, "WS" = would-strengthen.

# Sev Issue (deduped) Section / line The fix
1 MF "Structural / intrinsic" 5/6 ceiling over-claims from a null result over a finite grid; title sells the negative, and §5(iii) proposes a known way to break it — having it both ways. (R2-02, np-01, rig-02) Title L30; Abstract L42; §4.4 L149; §5 L257 Demote to "apparent / empirical ceiling across the configs tested." Add the rule-of-three bound (0/856 → 95% UB ≈0.35%). Reframe robust ceiling as 4/6. See §2(d).
2 MF Abstract buries the contribution and is a 335-word wall with a comma-splice ending. (abstract-contribution-buried, abstract-wall-of-text) Abstract L42 Rewrite to land what-was-built + headline result in sentences 1–2; break the run-on; cut percentages that reappear in Results. See §2(a).
3 MF bSi → diatomgraz mechanism never stated, yet it is the load-bearing constraint for the entire diatomgraz / SO-destroys-recovery story. (DOM-01) §3 obs channels L79; §4.7(ii) L215 Add 2–3 sentences with the steady-state relation + Brzezinski 1985 / Tréguer & De La Rocha 2013 cites. See §2(e).
4 MF No confidence intervals anywhere; the AOI-ablation contrasts ride on sub-30% rates at n=40–80. (rig-03) Tables 1–3; §4.7 L217 Add Wilson 95% CIs, at minimum to Table 3 and the 38/40 and 16/80 contrasts. See §3.
5 MF "Five independent pieces of evidence" are not independent — items 1–4 nest/share seeds from the 856 pool. (rig-01) §5 L238–245 Change to "five lines of evidence, of which the 2-AOI ablation is the only one on disjoint seeds." See §3.
6 MF Circularity/proxy threat unconfronted: targets generated by full 39-tracer Darwin; inversion is a 5–7-tracer box. Cal-grade distance conflates proxy bias with identifiability failure. (R2-01) §3 / §5 Limitations L259 One paragraph: Cal-grade distance bounds proxy + identifiability error, not identifiability alone; cite the existing synthetic-truth validation (STATUS v0.x–v1.8) if it ran on the current kernel, else flag as required future work. See §5.
7 MF Related work omits ACE2S (Clark 2026) and Neural GCM (Kochkov 2024) — both absent from references.bib. Note: the brief mislabels Neural GCM as an "existing cite"; it is not in the bib. (np-02, np-03, ml-4, R2-05, DOM/missing-ace2s) §2 L61; references.bib Add both bib entries; expand the related-work clause into a 3-class taxonomy. See §2(b).
8 MF DINN architecture under-specified (activations, widths, depths, LayerNorm, consistency-penalty form) → not reproducible. (ml-1) §3 L69–77 Add a compact 2-network spec block. See §4 note.
9 MF Fig 4 title says "857 seeds" vs canonical 856 everywhere else. (fig4-seed-count-mismatch) make_figures.py L351 Regenerate driving the title from len(data["seeds"]); do not touch paper's 856.
10 WS Which network produced the headline numbers is never stated; "DINNDeep … higher-quality" contradicts the project's own "DINN baseline only / recovers fewer Cal-grade" rule (STATUS L94). (method-dinn-vs-dinndeep, R2-09) §3 L69; §4.1 L87 State headline sweep used the baseline/shared DINN; re-describe DINNDeep's role accurately.
11 WS "exact landing" overstated (median 0.044 vs Carroll, with 80% of seeds outside tolerance); and Carroll R_PICPOC mis-rounded 0.043 (canonical 0.04245→0.042). (rig-06, DOM-05) §4.7(i) L213; §5 L236 "the median lands near Carroll's value, though only 20% of seeds are within Cal-grade." Fix 0.043→0.042 in both places.
12 WS scav_rat ratio arithmetic: 6.0e-7/1.7e-7 = 3.5×, not "~4×". (DOM-04) §4.7(i) L213 "~3.5× below."
13 WS 0/856 is not a matched control for the 2-AOI basin (all 856 are 3-AOI by construction). (rig-05) §4.7(iii) L217 Lead the contrast with the matched 3-AOI baseline 0/40, demote 0/856 to context.
14 WS Climatology confound not connected to the gradient — time-mean removes the seasonal signal that most constrains growth rates + R_PICPOC. (R2-07) §4.2 L114 / §5 L259 One sentence: the Hard/Partial tiers may partly reflect the climatology choice; time-resolved fitting (Next Steps) is the discriminating test.
15 WS Per-parameter gradient is pooled over a non-random adversarial sweep — present as config-conditional, not an intrinsic obs-set property. (rig-09) §4.2 L114 One caveat sentence; reconcile with the "configuration-dependent" finding.
16 WS No figure for the AOI ablation — the most novel finding is a dense 6×4 table. (no-aoi-ablation-figure, figure-count-vs-page-budget) §4.7 / Figures Add a 6×4 heatmap; cut/inline Fig 5 (three scalars) to hold net figure count. See §4.
17 WS Novelty vs BINN stated only as a domain port. (ml-3, np-04, R2-03) §2 L61 Add a Contributions clause: multi-region joint loss + per-region attribution + recoverability/ceiling diagnostic are new relative to BINN. See §2(b).
18 WS raissi2019pinn orphaned (in bib, never cited; won't render). (ml-10, np-07) §3 L75; bib Cite as inverse-problem lineage — "PINN-style inverse / hard-physics variant," not "we are a PINN."
19 WS Mutex "binary, regardless of magnitude" over-deterministic at 8×(0/10); "complementary" 5/6 rests on two single seeds; closeout itself leans "n=10 noise." (rig-07, rig-04) §4.6 L182; §4.4 L159 Soften to "every config tested (8, down to 0.02) → 0/10, no detectable dose-response"; label complementarity anecdotal; cite closeout's own noise caveat.
20 WS "Carroll-N" vs "Carroll-6" used interchangeably for a fixed set of six (even within one sentence). (R2-10) Throughout (8× Carroll-N) Standardize on "Carroll-6."

Minor/cosmetic (batch in one pass): R_PICPOC sigmoid bound is linear over ~100× range — note as candidate log-bound fix in Next Steps (ml-9); structural-identifiability caveat on fixed R_SI_C/W_SINK (DOM-08); "near-steady state" tracer-resolved qualifier + report Wave-4 N_STEPS=1000 outcome (DOM-09, ml-6); ±40% Cal-grade threshold motivation (DOM-12); air-sea CO₂ flux is model-output target vs independent obs (DOM-11); Table 3 all-tied alpfe-row bold convention (table3-bold-convention); "to our knowledge, the first…" hedge (R2-11); §4 roadmap sentence (mutex-section-placement); Carroll-2020-vs-2022 caption note (carroll2020-vs-2022-optimum-label).


2. Drafted replacement prose (LaTeX-ready)

(a) Tightened abstract — lands the contribution in the first two sentences

Replace the entire abstract body (L42). This is ~230 words, leads with what-was-built + headline result, breaks the comma-splice, de-bolds the inline numbers, and phrases the ceiling as "below 6/6" (the grounded fact) rather than a stable 5/6.

\noindent We present DarwinDiff, a differentiable PyTorch reimplementation of
ECCO-Darwin's Darwin biogeochemistry that recovers the six Green's-functions-calibrated
``Carroll-6'' parameters by gradient descent through the box model: a small per-cell
network predicts the parameters, sigmoid-bounded into Carroll's published ranges, and
a z-scored observation loss backpropagates end-to-end through a forward-integrated
source-minus-sink (SMS) kernel. Across an 856-seed / 86-configuration sweep plus a
200-seed area-of-interest (AOI) ablation, the iron pair (\texttt{alpfe}, \texttt{scav\_rat})
recovers reproducibly at 38/40 (95\%, Wilson 95\% CI 83--99\%) Cal-grade --- to our
knowledge the first reproducible differentiable parameter-recovery result for ECCO-Darwin
biogeochemistry. Three findings follow. \textbf{(1)} A clear recoverability gradient
sorts the six parameters into reliable (the iron pair), partial (the two growth rates),
and rarely-recovered (\texttt{diatomgraz}, \texttt{R\_PICPOC}) tiers; the typical seed
recovers two or three of six. \textbf{(2)} This gradient is configuration-dependent, not
intrinsic: a 2-AOI ablation that drops the Southern Ocean recovers \texttt{R\_PICPOC}
(20\% vs.\ 0\% in the matched 3-AOI baseline) and \texttt{diatomgraz} (85\% vs.\ 0\%),
because the 3-AOI loss landscape was actively destroying them via inter-regime tension.
\textbf{(3)} No seed reaches 6/6 (0/856; 95\% upper bound on the true rate $\approx$0.35\%),
and the reproducible ceiling is 4/6 --- 5/6 occurred twice, both single-seed and
unreproduced. We attribute this apparent ceiling to parameter conservation under
$\sim$5 effective observational constraints on six parameters; we do not claim it is
intrinsic, and a per-AOI-gating architecture (Discussion) may lift it. DarwinDiff is
therefore a \emph{partial}, honest replacement of Green's-functions calibration.

Replace the third Background paragraph (L61). This adds the two missing SOTA cites, draws the three-class taxonomy, and states the methodological delta over BINN explicitly. Requires adding clark2026ace2s and kochkov2024neuralgcm to references.bib (entries below) and citing raissi2019pinn.

The methodological precedent for our approach is BINN, the Biogeochemistry-Informed
Neural Network of \citet{xu2025binn}: a small MLP predicts soil-carbon parameters per
location, sigmoid-bounded into prior ranges, fed into a differentiable PyTorch
reimplementation of CLM5 with an SOC-profile loss that backpropagates end-to-end.
DarwinDiff is a PINN-style inverse problem in the sense of \citet{raissi2019pinn} ---
fitting observations by differentiating through a forward integrator --- but with the
physics residual hard-coded as an exact port of the Darwin SMS kernel rather than a soft
PDE-residual penalty over a learned field. Three things distinguish it from a domain port
of BINN: (i) the target is a data-assimilative, operational ocean model with
independently-published Green's-functions optima, so recovery is scored against an
external calibration rather than self-consistency; (ii) we extend the single-site BINN
loss to a multi-region joint loss whose per-region parameter attribution is, we show, the
operative recoverability knob; and (iii) our central contribution is an identifiability
analysis --- a per-parameter recoverability gradient, a structural ceiling, and AOI
attribution --- that BINN does not attempt.

A parallel and now-dominant line of work builds neural Earth-system emulators for forward
prediction: Neural GCM \citep{kochkov2024neuralgcm}, ACE2S \citep{clark2026ace2s}, and,
within ocean BGC, Neural-BGC \citep{oualalachkar2026neuralbgc} (pure-NN emulation of
O$_2$ and NO$_3$); differentiable Veros \citep{meunier2025veros} makes a dynamical core
differentiable but targets the forward model. These learn or accelerate dynamics and
discard the model's interpretable parameters. DarwinDiff instead solves the inverse
problem --- its output is the six mechanistically named Carroll-6 values. Emulators answer
``what happens next''; we answer ``what is the parameter value.''

New bib entries:

@article{kochkov2024neuralgcm,
  author  = {Kochkov, Dmitrii and Yuval, Janni and Langmore, Ian and Norgaard, Peter and others},
  title   = {Neural General Circulation Models for Weather and Climate},
  journal = {Nature},
  volume  = {632},
  pages   = {1060--1066},
  year    = {2024},
  doi     = {10.1038/s41586-024-07744-y},
}

@article{clark2026ace2s,
  author  = {Clark, Spencer K. and others},
  title   = {{ACE2S}: A Stochastic Neural Earth-System Emulator},
  journal = {arXiv preprint arXiv:2606.07928},
  year    = {2026},
  url     = {https://arxiv.org/abs/2606.07928},
}

Flag for Lucas/Jon: verify the ACE2S author list and exact title before submission — I have the arXiv id (2606.07928) from the brief but should not fabricate the full author string. The Kochkov Nature DOI is correct.

(c) The single strongest HONEST headline sentence

For the abstract / talk / any one-liner. Honest about "partial," signals generality beyond one model, avoids an undefendable category-level "first":

DarwinDiff recovers a subset of an operational ocean-biogeochemistry model's Green's-functions-calibrated parameters by end-to-end differentiation — the iron pair reproducibly (38/40), four of six under at least one observation configuration, but never all six at once — establishing differentiable parameter recovery as a partial, honest alternative to Green's-functions calibration for ECCO-Darwin.

(d) Title + the over-claiming framing fixes

Title (L30) — drop the bare "Structural":

\title{\textbf{Differentiable Parameter Recovery for ECCO-Darwin Biogeochemistry:\\
A BINN-Style Framework and an Apparent 5/6 Recoverability Ceiling}\\[0.4em]
\large\textit{Internal white paper --- DarwinDiff project}}

§4.4 (L149) — add the rule-of-three bound and the 4/6-is-the-robust-ceiling statement:

The headline empirical finding: \textbf{0/856 seeds reach all six Carroll-6 parameters
jointly at Cal-grade} (one-sided 95\% upper bound on the true 6/6 rate $\approx$0.35\%).
The reproducible ceiling is 4/6 (69 seeds, 8.1\%); 5/6 is the rare observed maximum, twice
across 856 seeds, and did not reproduce at $n{=}20$.

§5 closing (L257) — reconcile the have-it-both-ways tension:

\textbf{What this is \emph{not} yet:} a full replacement of Carroll's six-parameter
calibration. We recover two reproducibly under the 3-AOI F2 configuration and four under
at least one AOI configuration, but never all six in a single configuration. This ceiling
is \emph{apparent} --- empirically unbroken across every lever, AOI mix, and composition we
tested, and consistent with parameter conservation under $\sim$5 effective constraints on
six parameters --- but we do not claim it is intrinsic: the per-AOI-gating architecture of
(iii) is a concrete, untested route that may lift it.

(e) bSi → diatomgraz mechanism (must-fix DOM-01)

Append to the §3 "Observation channels" paragraph (L79); the derivation is already in src/darwindiff/silica.py:

Surface bSi enters as a steady-state diagnostic, $\mathrm{bSi}_1 = R_{\mathrm{Si:C}}\,
(\mathrm{mort}_{\mathrm{diatom}} + \mathrm{graze}_{\mathrm{diatom}})/W_{\mathrm{sink}}$,
where the grazing term $\mathrm{graze}_{\mathrm{diatom}} = g_{\mathrm{diatom}}\,
G_0\, P_{\mathrm{diatom}}$ carries \texttt{diatomgraz} directly. Anchoring \emph{absolute}
bSi therefore pins $P_{\mathrm{diatom}}\cdot g_{\mathrm{diatom}}$; because the Chl1
z-score loss quasi-pins diatom biomass $P_{\mathrm{diatom}}$, the grazing rate
$g_{\mathrm{diatom}}$ (i.e.\ \texttt{diatomgraz}) becomes identifiable. We hold
$R_{\mathrm{Si:C}}=0.13$ \citep{brzezinski1985} and the subsurface dissolution rate at
$0.015\,\mathrm{d}^{-1}$ \citep{treguer2013}; both are fixed literature point values, so
absolute \texttt{diatomgraz} recoveries are conditional on them (Limitations).

Then at §4.7(ii) (L215), change "surface biogenic-silica observations in the Southern Ocean prefer a low \texttt{diatomgraz} value" to "via the steady-state bSi diagnostic of §3, surface biogenic-silica observations in the Southern Ocean prefer a low \texttt{diatomgraz} value." Add brzezinski1985 and treguer2013 to the bib.


3. Rigor tightenings

Wilson 95% CIs — add to these specific numbers (I computed them; safe to drop in):

Claim Point Wilson 95% CI Where
Iron pair F2 38/40 = 95% [83.5%, 98.6%] Table 1 L106, Fig 3 caption L143
eqp+natl R_PICPOC 16/80 = 20% [12.7%, 30.0%] §4.7(iii) L217
eqp+natl k≥4 24% (≈19/80) [15.8%, 34.1%] Table 3 L205
alpfe (n=856) 84% [81.4%, 86.3%] Table 2 L124
R_PICPOC (n=856) 3% [2.1%, 4.4%] Table 2 L129
3-AOI matched (joint R_PICPOC+diatomgraz) 0/40 one-sided UB 8.8% §4.7(iii)

Implement as a one-line table note ("proportions reported with Wilson score 95% CIs") plus parenthetical CIs in Table 3 and the 38/40 and 16/80 contrasts. The n=856 CIs are narrow and strengthen the gradient — include them.

The "five independent pieces of evidence" — neutralize the non-independence critique (rig-01). Items 1–4 share/nest seeds from the 856 pool (the Wave-6 config, both 5/6 base configs, and the n=20 extension are all inside the 856); only item 5 (the n=80 eqp+natl ablation) is on disjoint seeds. Rewrite the lead-in (L238) and item framing:

The \textbf{5/6 ceiling} is an empirical upper bound holding across every AOI
configuration we tested. Five lines of evidence support it --- of which the 2-AOI
ablation (item 5) is the only one computed on seeds disjoint from the 856-seed pool;
items 1--4 are nested or overlapping tallies from the same sweep, so they corroborate
rather than independently confirm:

Then soften item 4's mean-k claim (rig-08) — rest it on the categorical outcome ("0/10 at 5/6; iron pair held 9/10; neither R_PICPOC nor diatomgraz recovered"), not on the n=10 mean-k ordering (2.00 vs 2.40/2.70 carries no dispersion). And add to §4.4 (L159) that at n=20 each, a 1/20 rate has Wilson CI up to ~24%, so the two 5/6 events' "complementarity" is anecdotal, aligning with the closeout's own "Wave 5 leans toward statistical noise at n=10."

Mutex determinism (rig-07). Replace "regardless of magnitude … binary, not graduated" (L182) with: "every PIC-bearing configuration tested (8 configs, doses down to 0.02) drove iron-pair recovery to 0/10, with no detectable dose-response; we interpret this as an effectively binary basin flip — the nulls clear the 40% tolerance at every dose but do not prove a literal step function." Keep the mechanism as a hypothesis. Define M_tot at first use (DOM-06): it is total column biomass/production; trace the budget (PIC anchor fixes production-scaled PIC → total production → under iron limitation the iron budget → alpfe/scav_rat compensate) or soften "parameter conservation" to "a tight degeneracy ridge."

Numeric corrections (one pass): §4.7(i) L213: "~4×" → "~3.5×" (6.0e-7/1.7e-7 = 3.53); "exact landing" → "the median lands near Carroll's value, though only 20% of seeds are within Cal-grade tolerance"; Carroll R_PICPOC "0.043" → "0.042" here and at L236 (canonical 0.04245).


4. Figures and tables

  • Cut/inline Fig 5 (Wave-6 three-bar chart). It renders three scalars (2.00 vs 2.40/2.70) the sentence at L170 already states in full — lowest information-per-figure in the paper.
  • Add an AOI-ablation figure to §4.7: a 6×4 heatmap (Carroll-6 rows × {3-AOI, eqp+natl, eqp+SO, eqp-only} columns; cell color = Cal-grade rate, percentage printed in-cell). Highlight the eqp+natl column. This visualizes the paper's most novel finding (the diatomgraz 100→85→15→0% monotone drop and the R_PICPOC/diatomgraz basin). All data is in aggregated_v3.1.json — no new computation. Net figure count stays at 5.
  • Keep Fig 4 — distribution shape (R_PICPOC sitting above unity) is information Table 2's rates cannot convey; the finding that wanted it cut self-corrected.
  • Fig 4 title: regenerate from make_figures.py so the title reads 856 (drive it from len(data["seeds"]), don't hardcode).
  • Table 3 bold: the all-tied 100% alpfe row makes "bold = highest per row" vacuous — drop the bold on that row or add "ties left unbolded." (The summary-row claim in the original finding is a misread; those rows are already consistent.)

5. Reviewer-2 defense — the 4 attacks most likely to land

Attack 1 — "You inverted a different dynamical system than the one that made your targets, so Cal-grade distance isn't recovery skill." (The single biggest threat; R2-01.) Pre-empt with a Limitations paragraph:

\paragraph{Proxy versus identifiability.} The fit targets are full-Darwin (39-tracer,
advection-coupled) v05 output, while the inverted model is a per-cell 5--7-tracer box.
Cal-grade distance from Carroll therefore bounds the \emph{sum} of structural proxy error
and identifiability error, not identifiability alone; a ``miss'' (e.g.\ \texttt{R\_PICPOC}
at 3\%) cannot yet be cleanly attributed to non-identifiability rather than box-vs-full
mismatch. A synthetic self-consistency control --- generating targets by running the same
box model forward at Carroll's values, then re-running recovery --- would separate the two:
if the gradient reproduces on self-generated targets it is an identifiability property of
the box model; if it dissolves, the ``Hard'' tier is largely proxy bias.

Then state whether the existing synthetic-truth validation (STATUS v0.x–v1.8) ran on the current 5-PFT/carbonate kernel. If yes, cite it as the control. If it predates the kernel extension, say so and list the re-run as required future work. (Do not claim it as done unless verified — the box was extended over time.)

Attack 2 — "Your headline is the negative result, but the negative result is a finite-grid null you yourselves propose to break." Defended by §2(d): demote "structural→apparent," add the rule-of-three bound, and reconcile with the per-AOI-gating proposal in one sentence.

Attack 3 — "The AOI 'reproducibility' cuts against you: the same scav_rat goes 95%→1%, diatomgraz 0%→85%, just by changing regions — these aren't recoveries of a physical optimum, they're loss-weighting artifacts." (R2-04.) The paper already frames this as "configuration-dependent," but add an independent identifiability cross-check sentence in §4.7: report the magnitude of the swings as a quantitative weak-identifiability signal, and add the finite-difference loss sensitivity at Carroll's optimum (cheap via autograd) as the diagnostic that grounds the gradient in loss geometry rather than optimizer outcomes. One sentence committing to it as the natural confirmatory test is enough if the computation is out of scope.

Attack 4 — "Why is recovering 'the easy two' (iron) novel, and what's new beyond BINN?" (R2-03.) Defended by §2(b) (Contributions clause) plus one sentence stating the iron pair is non-trivial because it needs the absolute-units GEOTRACES anchor — a z-scored pattern loss alone does not pin alpfe/scav_rat (project history). Reframe Discussion (i) from "fast 7-minute calibration" (engineering) toward "which parameters are constrained by current v05 obs" (science).


6. Submission readiness — internal white paper → submittable short/workshop paper

The doc-quality fixes above (§1–§5) make it a strong internal artifact. To clear an external bar (e.g., a NeurIPS/ICLR ML-for-physical-sciences workshop or a JAMES short communication), three things are load-bearing and go beyond cleanup:

  1. One independent validation experiment (the decisive one for Attacks 1 & 3). Either the synthetic self-consistency control (run the box forward at Carroll's values, recover, check the gradient reproduces) or the finite-difference / autograd Jacobian eigenspectrum at Carroll's optimum (6×6 Gauss-Newton on the bounded parameters — near-zero cost given the pipeline is already differentiable). A ~5-of-6 effective-rank result would convert "structural ceiling" from a null-result anecdote into a measured identifiability bound — the single biggest credibility upgrade available, and it directly earns back the word "structural." This is the one item I would not ship externally without.

  2. Architecture reproducibility block (ml-1): the two-network spec (DINN = Conv1×1(c→16)–Tanh–Conv–Tanh–Conv(→6); DINNDeep = pre-norm residual blocks, GELU, per-cell LayerNorm-over-channels) plus the consistency-penalty form. External reviewers will not accept "~454 weights, SST input only" as a network description. Note: STATUS.md says ~454 but the audit (ml-2) computes 406 (SST-only)/438 (default) — reconcile the weight count against the actual networks.py instantiation before quoting it.

  3. CIs + the non-independence honesty fixes (§3) are table stakes for a stats-literate venue; they are cheap and already drafted.

Everything else in §1 is polish that raises quality but is not a gate. The proxy-circularity paragraph (§5) is mandatory for either internal or external — it is the most obvious thing Jon (or any oceanographer) will ask, and the paper is currently silent on it.

Net assessment: the science is honest and the headline numbers are solid. The work to do is almost entirely defensive framing and rigor presentation, not new results — with the single exception of one identifiability/self-consistency control, which is both cheap and the highest-value addition. I recommend the must-fix list (§1 items 1–9) plus the proxy paragraph before this goes any further, and the validation experiment before any external submission.

Key files referenced (all absolute): - Manuscript: C:\Users\Frank\OneDrive\Desktop\Github\ecco-darwindiff\docs\paper\main.tex - Bib: C:\Users\Frank\OneDrive\Desktop\Github\ecco-darwindiff\docs\paper\references.bib - Canonical data: C:\Users\Frank\OneDrive\Desktop\Github\ecco-darwindiff\docs\paper\figures\aggregated_v3.1.json - Closeout: C:\Users\Frank\OneDrive\Desktop\Github\ecco-darwindiff\docs\findings\v3.1_closeout.md - bSi derivation (for DOM-01): C:\Users\Frank\OneDrive\Desktop\Github\ecco-darwindiff\src\darwindiff\silica.py - DINN rule (for method-dinn fix): C:\Users\Frank\OneDrive\Desktop\Github\ecco-darwindiff\STATUS.md (L66, L94, L107) - Fig 4 title fix: C:\Users\Frank\OneDrive\Desktop\Github\ecco-darwindiff\docs\paper\figures\make_figures.py (L351)