Can the models that write the news also tell you what happens next — and say how sure they are? We took the 45 June 2026 events Pollar covered most broadly at birth, had a drafting model write binary forecasting questions about how each story would develop, struck the weak ones in a logged human vetting pass, and asked four models — Claude Fable 5, Pollar's production writer, Kimi K2.6, and GLM-4.6 — to put probabilities on the 28 surviving questions from 22 events, seen exactly as the stories stood at birth and again at each story's midpoint. Every question was resolved by a forecast-blind resolver; the 20 close-call or provisional-VOID cases received a context-isolated second blind review by another instance of the same model family. The headline: on these 28 retrospective questions, the preregistered paired mean Brier difference, Fable 5 minus the production writer at the birth snapshot, was −0.154 (95% CI [−0.254, −0.072], event-clustered bootstrap; negative favors Fable). Separately and descriptively, the production writer's mean Brier score was 0.276 versus 0.250 for the analytic 0.5 reference. Everything beyond the single preregistered contrast — calibration curves, Murphy components, the open-model table, update behavior — is descriptive, and labeled so. This estimate describes one preregistered paired contrast on 28 outcome-aware retrospective questions from 22 events; it is not an estimate of prospective forecasting performance. The preregistration survived six external audit rounds; 27 deviations are logged; every raw artifact is linked below.
A forecast without an error bar is a guess wearing a suit. Run 1 exists to build and pressure-test the machinery — frozen protocol, blind resolution, clustered uncertainty — on resolved ground, where mistakes are cheap, before run 2 makes forward forecasts in real time, where they are not. With 28 questions from 22 events (kept by a vetter instructed not to inflate n), inference uses event-clustered intervals, and we expected them too wide to separate frontier models: for external context, on ForecastBench (Karger et al., forecastbench.org) the difficulty-adjusted Brier gap between GPT-4 and GPT-4.5 was roughly 0.03. The observed gap here is larger than that context suggested it would be; the sensitivity paragraph in the results section is where we tried to break it.
| arm | Brier (question-wtd) | 95% CI | Brier (event-wtd) | log score | n (birth) | failed calls |
|---|---|---|---|---|---|---|
| Claude Fable 5 | 0.122 | [0.046, 0.207] | 0.108 | 0.392 | 28 | 0/56 |
| Kimi K2.6 | 0.180 | [0.083, 0.290] | 0.185 | 0.544 | 25 | 5/56 |
| 0.5 reference | 0.250 | — | 0.250 | 0.693 | 28 | — |
| production writer | 0.276 | [0.157, 0.397] | 0.273 | 0.842 | 28 | 2/56 |
| GLM-4.6 | 0.289 | [0.173, 0.416] | 0.302 | 0.878 | 28 | 0/56 |
The one preregistered claim: paired per-question Brier, Fable 5 minus the production writer, birth snapshot: −0.154 [−0.254, −0.072], all 28 pairs complete-case (no drops, no VOIDs). Every other number on this page is descriptive.
Sensitivity (the preregistered attempts to break it): re-scoring failed calls at 0.5 leaves the headline identical (both primary arms answered all 28 questions at birth). Zero VOIDs were issued, so the preregistered VOID stress bound collapses to the point estimate. Five forecasts across the run landed exactly on 0 or 1 (one production, two Kimi, two GLM; none from Fable) and are reported unclipped in the results file. The full per-question record, horizon cross-tab, and question-flow counts ship in the artifact bundle.
Murphy decomposition (exploratory, unstable point diagnostics — the protocol states plainly that components at this n are not evidence of a behavioral trait): reliability / resolution / uncertainty were 0.028 / 0.146 / 0.230 for Fable and 0.081 / 0.030 / 0.230 for the production writer. These components describe this sample only.
Updating (birth → midpoint, exploratory; temporal change, not a causal effect of additional articles): all four arms had negative point estimates of midpoint score change — Fable −0.046 (n = 28, CI [−0.138, 0.031]), production −0.091 (n = 26, two calls dropped, CI [−0.210, 0.028]), Kimi −0.025 (n = 25, CI [−0.132, 0.075]), GLM −0.131 (n = 28, CI [−0.255, −0.025], the only interval excluding zero). 0.5-imputation variants are in the results file. Only 2 of 28 questions saw no new articles by midpoint; that stratum is too small to say anything.
Reliability diagram: five equal-count bins per primary-family arm, with fixed-membership bootstrap bands on the outcome rate, in the interactive figure; the open-tier arms are an opt-in exploratory overlay. No writer is called "calibrated"; n = 28 is a sketch, not a verdict.
Every question is a row: both snapshots' probabilities from every arm, the resolution with its evidence links, and the model rationales, expandable. Nothing is hidden — the flow counts reconcile all 178 drafts (90 wave-1, 88 wave-2) down to the scoreboard, including 9 drafting-time duplicates, one event that defeated the drafter twice, 56 vetting strikes by reason, and the 40 sensitive-screen exclusions.
We looked for market twins for all 28 questions on Polymarket, Metaculus, Kalshi, and Manifold under strict matching (same subject, same threshold, deadline within ±24h). Zero qualified; every near-miss and its failing axis is in the published match log (crowd-candidates-log.json in the bundle). Niche European news mostly has no liquid market twin; where our questions brushed real markets — Starmer's premiership, US–Iran, Spanish elections — the market's date buckets straddled our deadlines. An empty comparison table, honestly arrived at, beats a stretched one.
Preregistration v1.3, frozen through six external audit rounds before any forecast call · 27 logged deviations, each with date, section, and reason · event manifest (2,852 rows) · sensitive-screen dual-review log · both drafting waves, including the struck one · vetting log with every strike reason and the pre-annotator's published recommendations · corpus manifests with per-article content hashes · harness ledger with every attempt · resolution log with dual-blind review, four verbatim reconciliations, and process-channel evidence · crowd-anchor near-miss log · analysis code at its frozen tag, 14/14 frozen tests. Cost: $2.65 in the committed harness ledger, plus a console session of 44 calls billed off-ledger at listed pricing — roughly $12 all in.
Download everything: preregistration v1.3 · deviations log (27) · event manifest · sensitive-screen review · both drafting waves · vetting log · questions · corpus manifest · harness ledger · resolution log (sha256, evidence) · crowd near-misses · results · contamination dossier (redacted) · audit rounds 1–5 · advisor review.
@techreport{pollarlabs005,
author = {Pollar Research},
title = {Tomorrow, with error bars},
year = {2026},
url = {https://labs.pollar.news/experiments/calibrated-forecasts/paper},
note = {Pollar Labs Experiment 005, run 1 (backtest)}
}