EXPERIMENT 005 · PAPER · v2

Tomorrow, with error bars

Pollar Research
resolved July 2026 · every artifact published · figures live in the playground · revised after advisor review (published)
STATUS: BACKTEST · forecasts generated after outcomes were known · primary-comparison contestants have a documented pre-June training cutoff or a pre-June stable-alias release (alias-to-weights immutability assumed, not verifiable; disclosed); any others appear only in the contamination-sensitivity arm · outcome-aware question selection may introduce hindsight bias · full resolution and vetting logs public.

Abstract

Can the models that write the news also tell you what happens next — and say how sure they are? We took the 45 June 2026 events Pollar covered most broadly at birth, had a drafting model write binary forecasting questions about how each story would develop, struck the weak ones in a logged human vetting pass, and asked four models — Claude Fable 5, Pollar's production writer, Kimi K2.6, and GLM-4.6 — to put probabilities on the 28 surviving questions from 22 events, seen exactly as the stories stood at birth and again at each story's midpoint. Every question was resolved by a forecast-blind resolver; the 20 close-call or provisional-VOID cases received a context-isolated second blind review by another instance of the same model family. The headline: on these 28 retrospective questions, the preregistered paired mean Brier difference, Fable 5 minus the production writer at the birth snapshot, was −0.154 (95% CI [−0.254, −0.072], event-clustered bootstrap; negative favors Fable). Separately and descriptively, the production writer's mean Brier score was 0.276 versus 0.250 for the analytic 0.5 reference. Everything beyond the single preregistered contrast — calibration curves, Murphy components, the open-model table, update behavior — is descriptive, and labeled so. This estimate describes one preregistered paired contrast on 28 outcome-aware retrospective questions from 22 events; it is not an estimate of prospective forecasting performance. The preregistration survived six external audit rounds; 27 deviations are logged; every raw artifact is linked below.

1Why a backtest, and why error bars

A forecast without an error bar is a guess wearing a suit. Run 1 exists to build and pressure-test the machinery — frozen protocol, blind resolution, clustered uncertainty — on resolved ground, where mistakes are cheap, before run 2 makes forward forecasts in real time, where they are not. With 28 questions from 22 events (kept by a vetter instructed not to inflate n), inference uses event-clustered intervals, and we expected them too wide to separate frontier models: for external context, on ForecastBench (Karger et al., forecastbench.org) the difficulty-adjusted Brier gap between GPT-4 and GPT-4.5 was roughly 0.03. The observed gap here is larger than that context suggested it would be; the sensitivity paragraph in the results section is where we tried to break it.

2What we did (the short version)

  1. Events. All 2,852 events Pollar ingested June 1–20 went into a published manifest; 1,667 were eligible (English, multi-source at birth, no timestamp anomalies). A keyword screen plus dual review excluded 40 casualty-focused stories. The run corpus is the mechanical top 45 by source diversity at birth + 1 hour — outcome-blind with respect to ingest timing, conditional on present-day event-article membership (a disclosed limitation).
  2. Questions. A drafting model wrote 2 binary questions per event from birth-snapshot articles only. Wave 1 produced 90 drafts; five duplicates were struck during drafting, and all 85 drafts reaching vetting were then struck: our frozen drafter prompt had omitted the protocol's own source-rule constraints, so no draft named a resolvable retrieval rule. The preregistration caught its own bug in public; the prompt was amended (logged), the wave published as struck, wave 2 drafted under a mechanical constraint gate. GPT-5.6 Sol — no arm in this race — pre-annotated both vetting rounds; the human vetter confirmed 28 keeps and 56 strikes, overriding nothing: "Confirm all, don't inflate n."
  3. Forecasts. Every contestant saw the identical packet — full question templates plus the same eight-article snapshot the drafter saw — at two instants per event. Fable 5 ran through a console channel (44 calls, prompts byte-identical to the harness by hash); the production writer, Kimi K2.6, and GLM-4.6 ran through a logged harness; a 0.5 reference is computed analytically. Failure rates are scored, not hidden: Kimi failed 5 of 56 question-calls and the production writer 2, all by exhausting token budgets on two stubborn packets; Fable and GLM failed none.
  4. Resolution. Every question resolved YES, NO, or VOID strictly per its own frozen template by a forecast-blind resolver; the 20 close calls and provisional VOIDs were independently re-resolved by a second blind instance that never saw the first verdict. Four disagreements went to documented reconciliation — two were access-artifact reversals caught by positive evidence, and in the two bot-walled cases a forecast-exposed coordinator relayed named-source content through the preregistered process channel while the blind reviewers retained every YES/NO/VOID determination. Final: 10 YES, 18 NO, 0 VOID; the log, with every reconciliation reply verbatim, was hashed and committed before scoring ran.
  5. Scoring. Analysis code frozen and golden-tested before resolution; bootstrap seed and binning preregistered; the scoring run is one command against the committed log.

3Results

armBrier (question-wtd)95% CIBrier (event-wtd)log scoren (birth)failed calls
Claude Fable 50.122[0.046, 0.207]0.1080.392280/56
Kimi K2.60.180[0.083, 0.290]0.1850.544255/56
0.5 reference0.2500.2500.69328
production writer0.276[0.157, 0.397]0.2730.842282/56
GLM-4.60.289[0.173, 0.416]0.3020.878280/56
Table 1 — descriptive scores, birth snapshot (question-weighted mean Brier, lower is better; event-clustered 95% CIs; log score secondary; n = questions with a parseable forecast).

The one preregistered claim: paired per-question Brier, Fable 5 minus the production writer, birth snapshot: −0.154 [−0.254, −0.072], all 28 pairs complete-case (no drops, no VOIDs). Every other number on this page is descriptive.

Sensitivity (the preregistered attempts to break it): re-scoring failed calls at 0.5 leaves the headline identical (both primary arms answered all 28 questions at birth). Zero VOIDs were issued, so the preregistered VOID stress bound collapses to the point estimate. Five forecasts across the run landed exactly on 0 or 1 (one production, two Kimi, two GLM; none from Fable) and are reported unclipped in the results file. The full per-question record, horizon cross-tab, and question-flow counts ship in the artifact bundle.

Murphy decomposition (exploratory, unstable point diagnostics — the protocol states plainly that components at this n are not evidence of a behavioral trait): reliability / resolution / uncertainty were 0.028 / 0.146 / 0.230 for Fable and 0.081 / 0.030 / 0.230 for the production writer. These components describe this sample only.

Updating (birth → midpoint, exploratory; temporal change, not a causal effect of additional articles): all four arms had negative point estimates of midpoint score change — Fable −0.046 (n = 28, CI [−0.138, 0.031]), production −0.091 (n = 26, two calls dropped, CI [−0.210, 0.028]), Kimi −0.025 (n = 25, CI [−0.132, 0.075]), GLM −0.131 (n = 28, CI [−0.255, −0.025], the only interval excluding zero). 0.5-imputation variants are in the results file. Only 2 of 28 questions saw no new articles by midpoint; that stratum is too small to say anything.

Reliability diagram: five equal-count bins per primary-family arm, with fixed-membership bootstrap bands on the outcome rate, in the interactive figure; the open-tier arms are an opt-in exploratory overlay. No writer is called "calibrated"; n = 28 is a sketch, not a verdict.

4The ledger

Every question is a row: both snapshots' probabilities from every arm, the resolution with its evidence links, and the model rationales, expandable. Nothing is hidden — the flow counts reconcile all 178 drafts (90 wave-1, 88 wave-2) down to the scoreboard, including 9 drafting-time duplicates, one event that defeated the drafter twice, 56 vetting strikes by reason, and the 40 sensitive-screen exclusions.

5What the crowd knew

We looked for market twins for all 28 questions on Polymarket, Metaculus, Kalshi, and Manifold under strict matching (same subject, same threshold, deadline within ±24h). Zero qualified; every near-miss and its failing axis is in the published match log (crowd-candidates-log.json in the bundle). Niche European news mostly has no liquid market twin; where our questions brushed real markets — Starmer's premiership, US–Iran, Spanish elections — the market's date buckets straddled our deadlines. An empty comparison table, honestly arrived at, beats a stretched one.

6Limitations, in plain sight

  1. This is a backtest with outcome-aware question construction. The drafter, the vetting pipeline, and every human involved knew how these stories ended. Containment mechanisms (birth-corpus-only drafting inputs, frozen rubric, published logs) reduce hindsight bias but cannot eliminate it. This is an estimate on an outcome-aware retrospective question set, not evidence of prospective forecasting performance.
  2. Run-corpus selection is outcome-blind for article ingest timing only, conditional on present-day event-article membership.
  3. Resolver and second reviewer are model instances of the same family; their independence is architectural, not institutional. Twenty of the 28 resolutions received the second blind review; in the two bot-wall reconciliations, a forecast-exposed coordinator relayed named-source content while the blind reviewers retained every determination. The published log lets you re-adjudicate everything.
  4. The production arm enters on stable-alias release-date evidence; alias-to-weights immutability is an assumption its provider does not let us verify, and the alias's routing suffix permits provider routing variance within the pinned policy. The production model is not named (the one redaction, disclosed); its arm is not independently reproducible.
  5. 28 clustered questions from 22 events. The headline interval is wide; the point estimate should be read with the interval, not instead of it.
  6. The console arm's model identity is operator-attested, and the selected console settings were not independently captured; the committed call-file hashes establish input integrity only (deviation 26).
  7. The resolution pass ran on July 11, two days before the preregistered date, under an operator constraint (deviation 27); every deadline and grace window had verifiably expired. No final V3 VOID was issued, so the observed VOID-sensitivity bound was unchanged — but the counterfactual effect of waiting until July 13 cannot be verified.

7The paper trail

Preregistration v1.3, frozen through six external audit rounds before any forecast call · 27 logged deviations, each with date, section, and reason · event manifest (2,852 rows) · sensitive-screen dual-review log · both drafting waves, including the struck one · vetting log with every strike reason and the pre-annotator's published recommendations · corpus manifests with per-article content hashes · harness ledger with every attempt · resolution log with dual-blind review, four verbatim reconciliations, and process-channel evidence · crowd-anchor near-miss log · analysis code at its frozen tag, 14/14 frozen tests. Cost: $2.65 in the committed harness ledger, plus a console session of 44 calls billed off-ledger at listed pricing — roughly $12 all in.

Download everything: preregistration v1.3 · deviations log (27) · event manifest · sensitive-screen review · both drafting waves · vetting log · questions · corpus manifest · harness ledger · resolution log (sha256, evidence) · crowd near-misses · results · contamination dossier (redacted) · audit rounds 15 · advisor review.

8Cite

@techreport{pollarlabs005,
  author = {Pollar Research},
  title  = {Tomorrow, with error bars},
  year   = {2026},
  url    = {https://labs.pollar.news/experiments/calibrated-forecasts/paper},
  note   = {Pollar Labs Experiment 005, run 1 (backtest)}
}