# 005 run 1 — deviations from PREREG (freeze tag: labs-005-prereg-v1)

| Date | Section | Deviation | Reason |
|---|---|---|---|
| 2026-07-11 | §3 (drafter output contract) | Generation-failure made ATOMIC: one retry; a second failure yields ZERO drafts for the event (no per-question salvage); shortfall counted in the question-flow diagram. The "fewer (or zero) drafts" ambiguity deleted. | round-3 audit R3-B01; pre-call amendment, no forecast calls made |
| 2026-07-11 | §4 (vetting rubric) | "Naming a more specific source" removed from TIGHTEN entirely. Criterion (a) must be satisfied by the source as the drafter named it; a draft whose named source cannot resolve the question is STRUCK `ambiguous-resolution`. TIGHTEN now permits only adding timezone/units and replacing a vague entity reference with the exact name already present in the birth corpus. | round-3 audit R3-B02; pre-call amendment, no forecast calls made |
| 2026-07-11 | §5.1 (question template) | YES/NO conditions redefined as exhaustive conditional on no VOID condition holding. VOID taxonomy frozen to exactly three (V1 source unavailable/discontinued at pass date; V2 premise dissolved; V3 indeterminate at pass date after grace expiry). primarySource/fallbackSource must each be a named, URL-addressable page or endpoint plus a one-line retrieval rule; graceRule must state an ISO-8601 expiry instant. | round-3 audit R3-B03; pre-call amendment, no forecast calls made |
| 2026-07-11 | §6 (arms table) + dossier internal section | Production arm identity pinned: the `WRITER_MODEL` constant in `apps/pipeline/src` (agenticSummarization.ts) at the repo commit of tag `labs-005-prereg-v1.1`, recorded verbatim with routing suffix and provider policy in the contamination dossier's NOT-FOR-PUBLICATION section. Harness must assert at startup that the runtime-resolved id string-equals the pin and abort otherwise (Plan-2 requirement). Public text stays redacted. | round-3 audit R3-B04; pre-call amendment, no forecast calls made |
| 2026-07-11 | §7 (parse rules) | Duplicate question id within one response = that question FAILS for that arm; ids not in the packet are ignored and logged; both rules must be covered by harness unit tests before any call (Plan-2 gate). | round-3 audit R3-B05; pre-call amendment, no forecast calls made |
| 2026-07-11 | §9 + §10.6 (resolution disagreements; VOID bound) | Disagreement rule pinned: one documented reconciliation; persistent disagreement → final status VOID (conservative), both positions logged. VOID sensitivity bound computed complete-case (pairs where both arms have parseable forecasts), with the 0.5-imputation variant reported separately. | round-3 audit R3-B06; pre-call amendment, no forecast calls made |
| 2026-07-11 | §10.3, §10.4, §10.5 | Reliability per-bin CI = event-cluster bootstrap of the per-bin mean outcome with bin membership fixed at the point estimate (no re-binning per replicate; stated simplification). Murphy components demoted to point diagnostics only, no CIs (bootstrap-interval promise withdrawn). Updating analysis complete-case on questions parseable at BOTH snapshots, 0.5-imputation sensitivity separate. | round-3 audit R3-B07; pre-call amendment, no forecast calls made |
| 2026-07-11 | `apps/pipeline/test/forecastEval/scoring.ts` (pinned analysis code) | Added `binOutcomeCIs` (per-bin outcome bootstrap, membership fixed) plus unit tests in `forecastScoring.test.ts` — the Plan-2 gate implementation for the §10.3 rule above. Post-tag edit to the pinned scoring library, logged per §0. | round-3 audit R3-B07 (implementation); pre-call amendment, no forecast calls made |
| 2026-07-11 | §11 (output set) | Added the outputs promised elsewhere, each with location: primary paired-contrast estimate + 95% CI (headline line above Table 1); Murphy point-components table (appendix); birth-date × horizon-band distribution table (methods appendix); question-flow diagram incl. generation-failure shortfalls and screen strikes; sensitivity block (VOID bounds, 0.5-imputation, unclipped endpoint forecasts). | round-3 audit R3-B08; pre-call amendment, no forecast calls made |
| 2026-07-11 | §0, §1.1 (DECISION-E), §1.4 + `event-manifest.json` + `eligibleEvents.ts` | DECISION-E re-pinned: tie-break = `birthArticleCount` (articles with created_at ≤ birth + 1h, outcome-blind) descending, then id ascending; present-day `articleCount` removed from the rule. Sensitive screen moved BEFORE ranking (all eligible events, birth-corpus material only, reviewers blind to rank); backfill rule deleted; estimand limitation restated. Manifest regenerated from prod with `birthArticleCount` and re-ranked: the new manifest is canonical; the v1 top-45 is SUPERSEDED pre-call (3 of 45 boundary events changed; v1 retained in git history as an audit artifact). | round-3 audit R3-M01; pre-call amendment, no forecast calls made |
| 2026-07-11 | design.md §5 (statusline); §10.1, §12 + dossier | Statusline now reads "primary-comparison contestants have a documented pre-June training cutoff or a verified pre-June immutable release". Release-date evidence described as "indirect but strong", not "proof", in the statusline, PREREG §6/§10.1/§12, dossier, and design.md §1/§3; noted that the immutability assumption is what the R3-B04 pin enforces. | round-3 audit R3-M02; pre-call amendment, no forecast calls made |
| 2026-07-11 | §13 (manual channel) | For every Fable console attempt Jakub saves the OpenRouter generation id + console timestamp; the harness fetches `GET /api/v1/generation?id=...` metadata (model, tokens, provider, created) into the cost ledger as independent verification; call-file inputs and raw outputs sha256-hashed in the corpus manifest. | round-3 audit R3-M03; pre-call amendment, no forecast calls made |
| 2026-07-11 | Header | Stale "Status: DRAFT … pending the freeze tag" line removed; header now states v1.1, amended per the round-3 audit, all amendments pre-call and logged; tag `labs-005-prereg-v1.1`. | round-3 audit R3-m01; pre-call amendment, no forecast calls made |
| 2026-07-11 | §13 (budget) + pilot-report.md | Raw harness cost recomputed without the 90 phantom drafter-forecast calls: 315 calls, raw ≈ $1.57 harness + $8.77 Fable console ≈ $10.3; total ≈ $10–16 with margins (was "$2.08 raw / $11–17"). | round-3 audit R3-m02; pre-call amendment, no forecast calls made |
| 2026-07-11 | §1.1 (DECISION-E), §11 + design.md decisions log | "Outcome-blind by construction" and "selection is outcome-blind" claims WITHDRAWN as overclaims. Restated: the ranking quantities filter on article INGEST TIMES, which are immutable, but event-to-article MEMBERSHIP is read from the present-day DB and is not provably the as-of-birth state (the §2 corpus limitation); selection is outcome-blind with respect to ingest timing only, conditional on present-day membership. Disclosed in the estimand statement and the §11 paper-limitations list. | round-4 audit R4-B01; pre-call amendment, no forecast calls made |
| 2026-07-11 | §6, §9, §11 + design.md §4 | Second reviewer changed from Jakub (forecast-exposed as the manual-console operator) to a FRESH Claude instance spawned with only the question packet and the named resolution sources — it never receives any forecast, ledger entry, or arm output (blind by construction). Jakub reduced to process disputes that never touch YES/NO/VOID determinations. Disagreement protocol unchanged (one documented reconciliation; persistent disagreement = VOID), now between two forecast-blind reviewers. Complete resolution log sha256-hashed and committed BEFORE any scoring executes (mechanical ordering gate). Same-family reviewer limitation disclosed in §9 and the §11 paper-limitations list. | round-4 audit R4-B02; pre-call amendment, no forecast calls made |
| 2026-07-11 | §5.1 (VOID taxonomy) | V3 narrowed from the discretionary "outcome is indeterminate at the resolution pass date after grace expiry" to an observable condition: neither the named primary source nor the named fallback source has published the specific datum the resolution criterion names, by the resolution pass date (checked and screenshot/URL-logged for both sources). Conditional-exhaustiveness sentence updated to reference the narrowed V3. | round-4 audit R4-M01; pre-call amendment, no forecast calls made |
| 2026-07-11 | §6, §10.1, §11, §12 + design.md §1/§3/§5 statusline + dossier header | Primary-family admission standard WEAKENED and DISCLOSED (supersedes the v1.1 "verified immutable release" wording, which overclaimed): documented pre-2026-06-01 training cutoff OR pre-2026-06-01 release of a stable model alias; alias-to-weights immutability is an ASSUMPTION the provider does not let us verify, disclosed as a residual contamination channel for the production arm; the `:floor` suffix permits provider routing variance within the pinned provider policy (disclosed). Added to the §11 paper-limitations list. The primary contrast is RETAINED under the weakened, disclosed standard. | round-4 audit R4-M02; pre-call amendment, no forecast calls made |
| 2026-07-11 | §13 (manual-channel provenance) | Generation metadata no longer claimed to verify "settings". Evidence roles stated plainly: generation metadata = identity/route/time/tokens; console-settings screenshot = selected sampling settings; call-file header = intended settings. Jakub captures ONE console-settings screenshot per calling SESSION (model, temperature, max_tokens, response_format as selected), saved under `docs/labs-forecasts-2026-07-11/calls/provenance/`. Console settings are not modified within a calling session; any settings change ends the session and requires a fresh screenshot before further attempts. | round-4 audit R4-M03; pre-call amendment, no forecast calls made |
| 2026-07-11 | design.md §1 (cutoff screen), §3 (open-tier arms, fallback), §5 (bundle note) | Three stale design.md passages aligned with the v1.2 two-branch admission standard, which they contradicted: (1) the §1 exclusion rule "a model whose cutoff cannot be documented as pre-June 1 is excluded" re-keyed to "a model satisfying neither branch of the admission standard"; (2) the §3 preregistered fallback trigger "if the production model's cutoff cannot be documented" re-keyed to "if the production model satisfies neither branch" — under the old wording the fallback would have FIRED for the production writer (cutoff undocumented) despite PREREG saying it does not; the amended text states plainly that the production writer is admitted via the release branch (stable alias 2026-04-24) and the fallback does not fire; (3) the absolute open-tier claim "neither can contain the June events" made conditional on the disclosed alias-immutability assumption; bundle note now names the admission evidence (release date; cutoff undocumented). | round-5 audit R5-M01; pre-call amendment, no forecast calls made |
| 2026-07-11 | §0 (freeze mechanics) | Freeze metadata aligned with the v1.3 header: protocol version string `005-run1-prereg-v1.2` → `005-run1-prereg-v1.3`; lineage extended with "v1.3 = the round-5 design.md alignment wave, tag `labs-005-prereg-v1.3`". Metadata-only; no operative rule changed. Tag `labs-005-prereg-v1.3` re-pointed to this snapshot. | round-6 audit R6-m01; pre-call amendment, no forecast calls made |
| 2026-07-11 | §3 (drafter prompt + output contract) | Systematic wave-1 vetting failure: the checkpoint-FINAL drafter prompt listed the §5.1 template FIELD NAMES but not the v1.1 field CONSTRAINTS (source = URL + one-line retrieval rule; graceRule = ISO-8601 expiry instant), so all 85 vetted drafts carried bare-URL sources (85/85, verified mechanically) and instant-less grace rules (85/85) and failed KEEP criterion (a) — first-failing-criterion reason `ambiguous-resolution` — under honest rubric application (advisory pre-annotation: GPT-5.6 Sol, non-contestant, disclosed; strikes confirmed by the human vetter). Pre-call remedy: prompt amended to state both constraints; a mechanical §5.1-constraint check added to the draft-time generation-failure definition (same atomic retry); drafting re-run as wave 2 under the amended prompt (fresh 1+1 budget per event, wave recorded per draft and per ledger line; wave-1 ledger entries excluded from wave-2 resume). Wave 1 remains published in drafts.json and the vetting log (struck in toto). No forecast call has been made. | wave-1 vetting outcome; pre-call amendment |
| 2026-07-11 | §1.2 (horizon comparison precision) | The §5.1 mechanical assertion caught q-8d9dba5d-1 with deadlineUtc 255 MILLISECONDS below birth+7d: events_all.created_at carries ms precision (birth ...T10:27:57.255Z) while the drafter echoed the {MIN_DEADLINE} placeholder truncated to whole seconds (...T10:27:57Z). Pinned pre-call: horizon-boundary comparisons (deadline within [birth+7d, min(birth+30d, cap)]) are evaluated at WHOLE-SECOND precision (both sides floored to seconds), matching the precision the drafter receives. The question — a vetter-confirmed KEEP — remains in the set; n stays 28/22. The alternative (striking a valid question over a 255 ms echo artifact) was rejected as substanceless. No forecast call has been made. | wave-2 §5.1 assertion finding; pre-call amendment |
| 2026-07-11 | §1.2, §5.1 (disclosures at the Task-6 review) | Two clarifications, no behavior change beyond what rows 22-23 logged: (1) row-23 whole-second precision applies at BOTH §1.2 call sites — the eligibility check (check51) and the horizon-band stratification (horizonBand); the band computation would otherwise have thrown on the same 255 ms edge. (2) The plan's "voidConditions ⊆ {V1,V2,V3} instantiations" is NOT mechanically checkable (drafts carry unlabeled free-text instantiations); the mechanical §5.1 gate asserts non-empty string array only, and taxonomy conformance rests on the vetting pass plus a manual review-time mapping of all 28 kept questions onto V1/V2/V3 (no out-of-taxonomy condition found; reviewer-verified). §9 resolution should treat void-condition taxonomy fit as human-adjudicated, not machine-guaranteed. | Task-6 review findings 2-3; pre-call disclosure |
| 2026-07-11 | §6 (concurrency) | Harness RUN concurrency changed MID-RUN from the pinned "sequential (one request in flight)" to bounded parallelism (≤16 requests in flight) for the remainder of the forecast calls, on operator instruction (Jakub). The first 35 completed forecast calls ran sequentially; the rest run pooled. Analytic irrelevance: calls are independent (no shared state, per-call sampling), per-call attempt ordering and the 3-request budget are unchanged, and ledger integrity is preserved (all ledger appends execute on the single JS thread; every attempt still carries its own timestamp). Both regimes are distinguishable in the ledger by timestamp overlap. | operator instruction, mid-run; disclosed |
| 2026-07-11 | §13 (manual-channel provenance), §11 (limitations) | Manual-channel generation-metadata verification WAIVED for run 1 by operator decision: the per-attempt OpenRouter generation-id capture, the GET /api/v1/generation metadata fetch, console timestamps, and the per-session settings screenshot (R3-M03 and R4-M03 controls) were not executed. Identity and settings evidence for the Fable console arm is therefore: the operator's attestation, recorded verbatim ("IT WAS claude-fable-5"), the intended-settings call-file headers (sha256-committed pre-call), and output-form consistency (44/44 responses in the pinned json_object forecasts shape). The executing agent challenged the waiver once and was overruled; the plan's fetchGenerationMeta.ts obligation is void for run 1. Disclosed as limitations item (vi). | operator decision; post-forecast disclosure |
| 2026-07-11 | §1.2/§9 (resolution pass date), score.ts validator | Resolution pass executed 2026-07-11 evening instead of the pinned "2026-07-13 or later", on operator instruction (model-access constraint: the session executing resolution, scoring, and the paper is Fable 5, whose access expires within hours; the operator directed that this session finish the run). The pin's closure rationale is SATISFIED and verified mechanically: all 28 kept questions have deadline ≤ 2026-07-10T23:59:00Z and every graceRule instant ≤ 2026-07-11T18:00:00Z — no question can change truth value after the pass. Residual risk disclosed: V3 is pass-date-relative ("neither named source has published the datum BY the pass date"); a source that would have published by 07-13 but not by tonight yields a V3 VOID that the pinned date would have avoided; any V3 issued tonight is annotated with this caveat in the resolution log and counted in the VOID sensitivity bound. score.ts validateResolutionLog passDate bound edited 2026-07-13 → 2026-07-11T18:00:00Z (logged analysis-code edit; frozen scoring.ts untouched). | operator instruction (access expiry); disclosed pre-scoring |
