When two newsrooms publish two different numbers for the same event,
most readers never find out. This experiment renders the disagreement
itself. Every dispatch in one cluster is reduced to its claims, and the
claims are compared: corroborated, contradicted, compatible under
composition, silent, or stale. The specimen below is real. Over one
weekend our pipeline collected fifteen dispatches from fifteen newsrooms
in six languages, all counting the dead of one prison riot; the counts
disagreed for thirty hours. The claims were extracted by hand. The
instrument that would extract them automatically is the future
experiment; this page is its target interface.
1The artifact
Fig. 1 is the ledger: every source's latest claim on one
quantity, read as of a chosen minute. Fig. 2 is the same quantity over
thirty hours. Both are interactive; every value opens the verbatim
sentence it was extracted from, in its original language, with a link to
the dispatch.
EXPERIMENT 003 · REAL CLUSTER · CLAIMS EXTRACTED BY HAND, NOT BY THE PIPELINE
Fig. 1. The claim ledger. One real event, one quantity,
every source's latest claim as of a chosen minute. Struck values are
headlines contradicted by their own body text; Σ marks claims made
as a composition; ∅ means the source is silent on that quantity.
Click a row for the verbatim quote.
Fig. 2. The same quantity over thirty hours. Hollow
points are claims published after a higher official toll was already on
the record. The instrument's verdicts for notable pairs are listed
below. Click a point for the quote.
CONTRADICTSreuters.com vs ct24.ceskatelevize.cz · 25 versus 19, published 14 minutes apart, different sourcing chains: police plus hospital versus medical sources.
COMPATIBLEct24.ceskatelevize.cz vs aljazeera.com · Different surface numbers, same composition: 15 inmates plus 4 guards equals 19.
CONTRADICTSnoticiasaominuto.com vs itself · The headline says 15; the article's own lead counts 15 detainees plus 4 guards, which is 19.
CONTRADICTSlemonde.fr vs bbc.co.uk · 23 versus 25, 22 minutes apart, both citing authorities.
CONTRADICTSlemonde.fr vs bbc.co.uk · Six guards versus four guards among the dead.
CORROBORATEDspiegel.de vs theguardian.com · 26 dead with the same composition, seven guards, from two independent newsrooms.
CONTRADICTSspiegel.de vs tagesschau.de · More than 75 versus around 100 injured, both from German public-facing wires on the same day.
SILENTreuters.com · The highest early count never breaks down inmates versus guards.
STALEsapo.pt vs 7sur7.be · 25 published 96 minutes after the toll had officially risen to 27.
COMPATIBLEdw.com vs itself · Vague lead, exact body: 'more than two dozen' resolves to 'at least 25' within the same dispatch.
2The specimen
Negombo prison, Sri Lanka, 6 and 7 July 2026. The riot
entered our pipeline at 08:31 UTC on Monday as nineteen dead, on the word
of a hospital director. Fourteen minutes after the Czech public
broadcaster repeated nineteen, Reuters carried twenty-five from two
police sources. Le Monde's twenty-three arrived mid-afternoon citing the
authorities, twenty-two minutes after the BBC's twenty-five cited the
same. By evening the composition itself disagreed: four, six, or seven
guards among the dead. The toll settled at twenty-seven on Tuesday
morning, and one newsroom published twenty-five an hour and a half
later. None of this is a scandal. It is what honest early coverage of a
chaotic event looks like, which is exactly why an instrument that reads
the disagreement is worth building.
Quotes are verbatim from our ingestion record, timestamps
are from the source articles as collected, and the full extraction table
is downloadable: extraction.json,
extraction.csv,
cluster-query.sql.
3Journal
2026-07 · why this is a mockup and not a demo
The ledger, the timeline, and the verdict vocabulary are settled;
the claims behind them were extracted by hand from fifteen dispatches.
What does not yet exist is the extractor: span-grounded claim
extraction across six languages, at cluster scale, with the
compositional arithmetic done by the system rather than by us. That
work lands with the autumn extraction pipeline, and the detection run
it enables will be preregistered before it touches data.
2026-07 · what the ledger does not say
No newsroom named on this page was wrong. Early casualty tolls
move because hospitals, police, and ministries count different things
at different hours; a headline can lag its own body text by an editing
cycle. The instrument's thesis is precisely that divergence is a
signal to inspect, not an accusation. Where a claim was published
after a higher official figure, the verdict is stale, and that is the
strongest word this page uses.
2026-07 · what would kill it
Extraction precision below usefulness on multilingual input, and
cost per cluster at frontier prices. A ledger that misreads one number
in twenty is worse than no ledger. Both get bench measures before
this experiment claims anything.
4Reported results
None yet. Under Rule 1, this
experiment cannot appear in an edition until it has a bench entry. The
first extraction-precision run publishes here, whatever it says.