# Experiment 005 spec — external audit, round 1

Auditor: GPT-5.6 Sol via Codex, 2026-07-11, read-only, run against spec v1.
Verdict: "sound with mandatory revisions, but not sound to run as written."
All findings accepted; spec v2 (design.md) is the response.

---

The protocol is **sound with mandatory revisions, but not sound to run as written**. Its paired-question structure, Brier scoring, public ledger, and artifact-first reporting are promising, yet the advertised "leakage-proof" claim is untenable because questions and criteria are selected after outcomes are known; the corpus is not demonstrably reconstructed as it existed at each snapshot; the population, exclusions, baselines, resolution rules, and primary estimand remain under-specified; and several analyses exceed what roughly 50 clustered questions can establish. Run 1 can credibly provide descriptive, event-clustered Brier comparisons and an exploratory updating analysis. It cannot establish precise calibration, stable Murphy components, fine model rankings, or model superiority over a small opportunistic crowd subset.

## Findings

| Severity | ID | Spec section | Exact quoted text (trimmed) | Category | Defect | Minimal correction |
|---|---|---|---|---|---|---|
| Blocker | B-01 | What this is; §5 | "a **leakage-proof backtest**"; "leakage-proof by training cutoff" | leakage | Training cutoff addresses possible contestant pretraining contamination, not hindsight in question construction, corpus reconstruction, resolution drafting, or anchor selection. | Rename it a "retrospective, cutoff-screened backtest" and state explicitly that outcome-aware design creates residual hindsight bias that cannot be eliminated in run 1. |
| Blocker | B-02 | §2 | "every draft is kept, tightened, or struck before any forecasting" | leakage | Outcome-aware humans can preferentially keep forecastable questions, strike hard ones, and tighten criteria around the observed result. | Freeze a detailed acceptance/edit/strike rubric before inspecting drafts, blind vetters to outcome material as far as practicable, prohibit outcome-responsive criterion edits, publish every original draft and edit diff, and acknowledge that only a prospective run can remove residual hindsight. |
| Blocker | B-03 | §1 | "Events born June 1–20… Horizons 7–30 days, so all questions resolve by ~July 10." | consistency | A June 20 event with a 30-day horizon resolves around July 20, so the dates do not support the claim that all outcomes are already known by July 10. | Either restrict each deadline to July 10 or later, change the event window, or delay resolution and forecasting until after the latest possible July 20 deadline. |
| Major | M-01 | §1; Open item 1 | "any model whose cutoff cannot be placed before June 1, 2026 carries an explicit caveat" | leakage | A caveat does not make an unverified or post-event model a valid member of a cutoff-screened comparison. | Exclude such a model from the confirmatory comparison or label its results as a separate potentially contaminated sensitivity analysis. |
| Major | M-02 | §2 | "A **non-contestant drafting model** (Gemini 3.5 Flash…) reads each birth-date snapshot" | leakage | "Non-contestant" does not prevent the drafting model from having learned the outcomes and using them to select unusually resolvable or salient questions. | Require and document a pre-June cutoff for the drafter, use a deterministic extraction rule, or label the generated question set as outcome-contaminated. |
| Major | M-03 | §1 | "two frozen article corpora on the event's articles"; "`pipeline_articles.created_at <= X`" | leakage | Filtering present-day event membership by article ingest time can backdate articles moved into an event by later merges, so it need not recreate what the pipeline associated with that event at the snapshot. | Reconstruct membership from timestamped association/merge logs, exclude affected merged events, or define the corpus honestly as a retrospective set rather than an as-of pipeline snapshot; publish its manifest and content hashes. |
| Major | M-04 | §1 | "birth date and calendar midpoint"; "`created_at <= X`" | protocol | Date-only cutoffs leave timezone, exact birth instant, inclusivity, midpoint rounding, and same-day post-birth articles ambiguous. | Define UTC timestamps from `events_all.created_at`, an exact horizon instant, arithmetic midpoint and rounding rule, and inclusive/exclusive SQL boundaries. |
| Major | M-05 | §2 | "Target: **~50 vetted questions** from ~80–100 drafts" | pre-registration | A target rather than a deterministic inclusion rule permits outcome-aware stopping or selection until an attractive set of 50 remains. | Score every draft passing a frozen eligibility rubric or predefine a deterministic sampling and tie-breaking rule independent of outcome and forecast values. |
| Major | M-06 | §2 | "Forecasters see only final vetted question text." | protocol | Forecasters may be scored against resolution details they never received, contrary to prediction-market practice. | Give every forecaster the identical frozen packet containing question text, YES/NO conditions, source hierarchy, deadline and timezone, fallback rules, and VOID conditions. |
| Major | M-07 | §3 | "a status-quo heuristic (probability that the no-change outcome holds)" | protocol | This specifies neither a numeric probability nor an outcome-blind rule for identifying which side is "no change." | Fix the probability, define a mechanical status-quo label before resolution, specify treatment of disputed cases, and publish the label for every question. |
| Major | M-08 | §3 | "where a question overlaps a historical Polymarket or Metaculus question"; "best effort, subset only" | pre-registration | Post hoc overlap judgments, source choice, timestamps, and illiquid-market selection can cherry-pick a favorable crowd comparison. | Freeze exact-match criteria, odds field, timestamp/timezone, source priority, liquidity threshold and missing-data rule; compare only on the identical matched subset and treat a small subset as case studies, not a benchmark ranking. |
| Major | M-09 | §4 | "resolves YES / NO / VOID by hand"; "VOIDs are excluded from scoring" | protocol | VOID is undefined and outcome-dependent exclusions can remove ambiguous or difficult questions after their behavior is known. | Predefine exhaustive YES, NO and narrowly limited VOID conditions; publish every void reason; report attrition by event and class; and include a sensitivity or bounding analysis. |
| Major | M-10 | §5 | "limitations (… one reviewer resolving…)" | protocol | A single unblinded resolver can see probabilities or model identities while deciding close cases, and no appeal rule is specified. | Blind the resolver to forecasts and identities, require independent second review for every VOID/dispute and preferably all questions, and publish disagreements; one resolver is acceptable only when criteria make resolution mechanical. |
| Major | M-11 | §4 | "~50 vetted questions… across ~40–50 events"; "**Rankings:** Brier score…" | statistics | Questions are clustered within events and weighted unequally while three writers, two snapshots, baselines, and several metrics create many unprioritized comparisons. | Declare one primary estimand and contrast, report both question- and event-weighted Brier scores, use paired event-cluster bootstrap intervals, and make remaining comparisons exploratory or apply a fixed multiplicity rule. |
| Major | M-12 | §4; Open item 4 | "binned calibration curve with confidence intervals"; "bin count and CI method chosen at analysis time" | statistics | Five approximate quintiles would contain only about ten questions each, while Wilson intervals assume independent observations and post hoc binning permits result-responsive presentation. | Freeze the binning rule and interval method before calls, show bin counts and mean predictions, use event-cluster uncertainty, and label the figure a descriptive reliability diagram rather than evidence that a writer is calibrated. |
| Major | M-13 | §4 | "Brier score with Murphy decomposition" | statistics | At this sample size the reliability and resolution components are highly bin-dependent and upward-biased by sampling noise, so they cannot sustain component-level model rankings. | Specify the exact decomposition and grouping algorithm now, bootstrap by event, and report components as unstable exploratory diagnostics with no claims of superiority. |
| Major | M-14 | §4 | "base-rate cowardice shows as low resolution" | presentation | This slogan presents a noisy decomposition term as a stable diagnosis and obscures that resolution depends on the sample, forecast grouping, and outcome mix. | Replace it with a technical definition and state that low empirical resolution at this n is not evidence of a general behavioral trait. |
| Major | M-15 | §4 | "log score secondary"; "probability in [0,1]" | statistics | A wrong forecast of exactly 0 or 1 has infinite log loss, but clipping, rejection, or infinity handling is unspecified. | Require probabilities inside fixed bounds such as [0.01,0.99], or preregister the exact clipping epsilon and show unclipped endpoint forecasts separately. |
| Major | M-16 | §4 | "movement, scored as movement toward or away from the resolved truth; did new evidence help" | statistics | Direction alone is not a defined proper score, and improvement can reflect shorter horizon, corpus reconstruction, or regression to the outcome rather than beneficial use of new evidence. | Define the primary update as paired change in Brier loss with a fixed sign convention, report event-cluster intervals and the no-new-article stratum, and call it midpoint score improvement rather than a causal effect of evidence. |
| Major | M-17 | §3–4 | "probability coherence — parseable, in range, rationale consistent with the stated number" | pre-registration | The design does not specify retries, parse failures, rationale disagreements, provider errors, duplicate attempts, or whether failed outputs enter denominators. | Freeze exact model IDs, prompts, parameters, retry limits, attempt-selection rule, and failure scoring; do not exclude an otherwise valid probability because a subjective rationale check fails. |
| Major | M-18 | §5 | "full artifact bundle"; listed files | presentation | The proposed bundle omits the exact prompts, model/version identifiers, request parameters, corpus manifests and hashes, attempt logs, scoring configuration, exclusions, and analysis code version needed to reproduce the result. | Add those items, plus an immutable preregistration snapshot and deviations log; if the production model remains unnamed, state that its result is not independently reproducible and publish auditable cutoff evidence. |
| Major | M-19 | §6 | "well under $10, input-token dominated" | consistency | This estimate conflicts with the existing bench, where 13 input-heavy calls cost $6.53 for Fable and $2.28 for Sol, implying tens of dollars when scaled to roughly 100 calls per writer before drafting and retries. | Price a representative birth/midpoint corpus through each exact model before approval and replace the promise with a budget range including retries. |
| Minor | m-01 | §2; §5 | "class tag (deal-closes, ceasefire-holds, record-stands, incumbent-leads, …)"; "filter by… class" | protocol | The open-ended taxonomy has no codebook, assignment rule, or minimum subgroup size. | Publish an exhaustive mutually exclusive codebook, assign classes before resolution, and declare classes descriptive-only in run 1. |
| Minor | m-02 | Decisions log | "New site status **BACKTEST** (CONCEPT → MOCKUP → BACKTEST → LIVE)" | consistency | The current house vocabulary is CONCEPT, MOCKUP and WORKING DEMO, so this is a replacement of the maturity grammar rather than merely an added status. | Either revise the method and home status key consistently or make BACKTEST an evidence label alongside the existing build-status vocabulary. |
| Minor | m-03 | §6 | "vetting pass ~1–2 h"; "resolution pass ~1 h" | protocol | The budget allows roughly one minute per draft or resolution, which is not credible for source hierarchy, ambiguity checks, edit logging, and independent adjudication. | Pilot ten questions, measure review time, and budget dual review for resolution and every exclusion. |

## Pre-registration checklist

- Freeze the immutable protocol, timestamp it, and define how deviations will be logged.
- Define event birth as an exact timestamp, multi-source eligibility, merge exclusions, horizon bounds, sensitive-event exclusion, and deterministic sampling/stopping.
- Reconcile the event window with the latest permissible resolution deadline.
- Define as-of corpus membership, timestamp boundary, timezone, article order, clipping, sorter/cap behavior, merge handling, and manifest/content hashes.
- Pin the drafting model, cutoff evidence, prompt, parameters, call count, deduplication and draft-generation rule.
- Freeze the keep/edit/strike rubric and publish every draft, version, reviewer decision and exclusion reason.
- Define the complete question template, including YES/NO, source hierarchy, deadline/timezone, corrections, postponements and VOID.
- Give forecasters the complete resolution packet, not question text alone.
- Pin contestant model IDs or immutable run identities, provider, prompts, parameters, randomness, concurrency and request timestamps.
- Define retries, provider failures, parse failures, endpoint probabilities, rationale mismatches and denominator treatment.
- Define the status-quo baseline's label rule and numeric probability without using outcomes.
- Define exact crowd-question matching, odds field, retrieval timestamp, liquidity threshold, source priority and missing-data treatment.
- Specify resolver blinding, second-review requirements, disagreement handling and the VOID sensitivity analysis.
- Declare the primary estimand, primary snapshot/contrast and whether weighting is by question or event.
- Define Brier and log-loss formulas, orientation, probability bounds or clipping, and tie handling.
- Choose the Murphy decomposition implementation, bin/grouping rule and bias/uncertainty treatment.
- Freeze calibration bins, displayed x-value, minimum bin count, interval level/method and event-cluster bootstrap settings.
- Define updating as a signed paired score change and specify treatment of questions receiving no new midpoint articles.
- Specify confidence intervals for aggregate scores and paired differences, bootstrap seed/replicates and multiplicity policy.
- Limit class and crowd-subset analyses to preregistered descriptive outputs unless minimum sample sizes are met.
- Freeze all tables, figures, filters and narrative headline rules so the most flattering metric cannot be selected afterward.
- Publish prompts, request configuration, corpus manifests, raw attempts, forecasts, resolutions, exclusions, scoring configuration, code revision, cost ledger and deviations log.
