EXPERIMENT 005 · TOMORROW, WITH ERROR BARS · BACKTEST

Tomorrow, with error bars

Pollar Research
resolved July 2026 · every artifact published · read the paper
STATUS: BACKTEST · forecasts generated after outcomes were known · primary-comparison contestants have a documented pre-June training cutoff or a pre-June stable-alias release (alias-to-weights immutability assumed, not verifiable; disclosed); any others appear only in the contamination-sensitivity arm · outcome-aware question selection may introduce hindsight bias · full resolution and vetting logs public.

Abstract

Can the models that write the news also tell you what happens next — and say how sure they are? We took the 45 June 2026 events Pollar covered most broadly at birth, had a drafting model write binary forecasting questions about how each story would develop, struck the weak ones in a logged human vetting pass, and asked four models — Claude Fable 5, Pollar's production writer, Kimi K2.6, and GLM-4.6 — to put probabilities on the 28 surviving questions from 22 events, seen exactly as the stories stood at birth and again at each story's midpoint. Every question was resolved by a forecast-blind resolver; the 20 close-call or provisional-VOID cases received a context-isolated second blind review by another instance of the same model family. The headline: on these 28 retrospective questions, the preregistered paired mean Brier difference, Fable 5 minus the production writer at the birth snapshot, was −0.154 (95% CI [−0.254, −0.072], event-clustered bootstrap; negative favors Fable). Separately and descriptively, the production writer's mean Brier score was 0.276 versus 0.250 for the analytic 0.5 reference. Everything beyond the single preregistered contrast — calibration curves, Murphy components, the open-model table, update behavior — is descriptive, and labeled so. This estimate describes one preregistered paired contrast on 28 outcome-aware retrospective questions from 22 events; it is not an estimate of prospective forecasting performance. The preregistration survived six external audit rounds; 27 deviations are logged; every raw artifact is linked below.

Loading the ledger and reliability diagram requires JavaScript.