# 005 paper review — academic advisor pass (GPT-5.6 Sol via Codex, 2026-07-12)

Verdict: MAJOR REVISIONS. All Major and Minor findings accepted and applied
in the same-date revision (paper v2 on the site), with one partial: the
ForecastBench context sentence was retained with an explicit external
citation and softened framing rather than deleted. Sol's three proposed
sentences were incorporated. Report verbatim below.

---

## 1. Advisor's verdict

**MAJOR REVISIONS.** The primary result is numerically sound, but the manuscript is not yet submission-worthy: it violates the preregistered headline discipline, overstates the resolution procedure, turns unstable exploratory diagnostics into behavioral claims, and omits several frozen outputs while claiming that a sensitivity block follows. The single most important sentence-level fix is paper-final.md:20: replace "forecast decisively better" and the attached production-versus-0.5 claim with a literal statement of the paired retrospective contrast, moving the latter into an explicitly descriptive sentence.

## 2. Line-referenced findings (abridged to finding + disposition; wordings applied in v2)

| Severity | Location | Finding | Disposition |
|---|---|---|---|
| Major | abstract ("decisively better" + prod-vs-0.5 inline) | Exceeds the registered estimand; exploratory claim not locally labeled | APPLIED: literal contrast statement + "Separately and descriptively" sentence |
| Major | abstract ("twice, independently") | False process summary — only 20/28 had second review; same-family reviewers | APPLIED: exact 20-in-scope wording |
| Major | "sensitivity block below" | No such block followed; frozen §11 sensitivity outputs missing from prose | APPLIED: sensitivity paragraph added to Results (imputation, VOID bound, endpoint forecasts, bundle pointer) |
| Major | Murphy ("stay honest… confidence was noise") | Converts unstable diagnostics into behavioral traits against the §10.4 label | APPLIED: point-values-only wording with the protocol's own caveat |
| Major | limitations ("capability signal under honest containment") | Generalizes beyond the contrast; self-certification | APPLIED: "estimate on an outcome-aware retrospective question set…" |
| Major | limitations item 6 | Hashes establish input integrity, not identity/settings; coordinator mediation and routing variance omitted | APPLIED: all three elements added (items 3, 4, 6) |
| Major | limitations item 7 ("risk never bound") | Zero VOIDs does not prove the July-13 counterfactual | APPLIED: counterfactual-unverifiable wording |
| Major | updating ("everyone improved") | Causal framing; selective reporting (GLM's was the only interval excluding zero) | APPLIED: all four arms with n, drops, CIs; temporal-change framing |
| Major | Table 1 | Missing n column; GLM event-weighted 0.303 should round 0.302461 → 0.302 | APPLIED: n column added; 0.302 |
| Minor | "all 85 drafts" wave-1 | 90 drafted, 5 duplicates, 85 vetted-and-struck | APPLIED |
| Minor | site metadata "28 events" | 28 questions from 22 events | APPLIED (card + abstract page) |
| Minor | "Why" section register + uncited ForecastBench 0.03 | Promotional register; uncited external comparison | PARTIALLY APPLIED: register softened, explicit citation added, "five times" removed; section retained |
| OK | BACKTEST statusline + Cite block | Consistent with frozen semantics | No change |

## 3. Number-trace outcome

All run-artifact numbers traced to results.json MATCHED except: GLM event-weighted rounding (fixed), missing Table-1 n (fixed). Items flagged MISMATCH solely because they live in artifacts other than results.json (manifest counts, console call count, deviation/audit counts, cost) were re-grounded in prose to their actual artifact sources; the cost line now separates the committed harness ledger ($2.65) from the off-ledger console estimate.

## 4. Advisor's proposed sentences

All three incorporated: the retrospective-estimate scope sentence (abstract + limitations 1), the 20-of-28 second-review disclosure with the coordinator's process-channel role (methods + limitations 3), and the early-pass counterfactual sentence (limitations 7).

