EXPERIMENT 004 · WHAT THE MEDIA IS OVERLOOKING

The naive gap table, and the instrument that replaces it

Pollar Research
opened July 2026 · last revised July 2026
STATUS: MOCKUP · real coverage matrix, synthetic readout · no bench entry yet

Abstract

"What the media is overlooking" is the easiest claim in media analysis to make and the hardest to earn. The naive version, a table of stories some cohort of sources did not cover, is a market-size chart in disguise: small cohorts cover less of everything. This page shows both halves of the honest version. The first half is real: one frozen week of our pipeline, 3,392 events, twelve provisional source cohorts, and the raw incidence matrix that any serious analysis must start from, shown without interpretation. The second half is deliberately synthetic: the readout the instrument will produce once absences are judged against a null that holds every outlet's activity and every event's reach fixed. Between the two halves sits a preregistration, and that is the point.

1The trap

In the week of 6 July our pipeline joined articles from roughly two hundred and thirty source domains to 3,392 events covered by at least three domains each. Group those domains by TLD and count distinct-domain coverage, and the temptation arrives immediately: sixty-three Polish domains touched 497 events, while seven Spanish domains touched 684. No editorial conclusion survives contact with the fourth column of Table 1.

EXPERIMENT 004 · REAL DATA, DESCRIPTIVE ONLY · NOTHING IN THIS MATRIX IS A FINDING
event (top 15 of 3,392 by source degree) .pl.com.de.fr.it.ch.gr.pt.be.es.ro.nl
Darline Graham sworn in as US senator after Lindsey…
Ann Widdecombe killing was 'targeted attack', UK co… ··
Bangkok pub fire death toll reaches 32 as investiga…
South Africa midfielder Jayden Adams found dead at… ·
IOC lifts suspension of Russian Olympic Committee,… ·
Macron's final Bastille Day parade showcases 500 co… ··
Senator's racist insults against Mbappé draw global…
Nolan's 'The Odyssey' opens to IMAX spectacle and a… ····
Typhoon Bavi weakens to tropical storm after slammi… ·
Bonnie Tyler, voice of 'Total Eclipse of the Heart'… ·
Apple sues OpenAI over alleged trade secret theft,… ·
Macron visits Fontainebleau as fire ravages 2,000 h…
New York Times asks court to quash subpoenas served… ·
Witnesses Contradict ICE Account of Fatal Houston S… ··
Rajoy's 'no Frenchmen' column on France's World Cup… ···
domains in cohort63541413119887776
events touched, of 3,3924971637568486407261399243256684324122
Fig. 1. One week of raw material. Coverage is a distinct source domain publishing at least one article joined to the event, never article volume. Rows are the window's fifteen highest-degree events that carry a published English title; the downloadable matrix has no such restriction. Columns are provisional source cohorts named by TLD. The gaps are visible and this page will not interpret them; judging which absences are surprising is the preregistered study's job, not the eye's. Click a row or a column header for the counts.
rankcohortevents toucheddomains doing the touching
1.com163754
2.es6847
3.de56814
4.pl49763
5.fr48613
6.it40711
7.gr3998
8.ro3247
9.ch2619
10.be2567
11.pt2438
12.nl1226
Table 1. The naive ranking, with the column that disarms it. Read without the fourth column, this table says Spanish sources out-cover Polish sources; with it, the table says our Spanish cohort is seven high-volume nationals while our Polish cohort is sixty-three mostly local outlets. A gap table that ignores this is a chart of publisher structure wearing the costume of editorial judgment.

2The readout

The corrected question is narrower and better: given everything this matrix already fixes, an outlet's weekly volume within its beat, an event's total reach, was this particular absence still unlikely? Fig. 2 is what answering it will look like. Every number in it is invented; publishing real surprises today, before the protocol freezes, would let the design see the answers it is supposed to be judged against.

EXPERIMENT 004 · SPECIMEN · SYNTHETIC DATA, FICTIONAL EVENTS · THE INTERFACE IS WHAT IS UNDER TEST
primary readout · excess rate of high-expectation omissions: +2.1 pp · randomization interval [-0.4, +4.8] · 2,000 degree-preserving draws · specimen numbers
eventcohortp₀ under the nullcomparable eventsverdict
Methane fee clears committee stage cohort II 0.94 214 OMITTED·SURPRISING
Rail corridor co-funding signed cohort I 0.88 178 COVERED
Deposit-return scheme postponed cohort III 0.92 203 COVERED
Grid interconnector tender annulled cohort II 0.83 190 OMITTED·SURPRISING
Border data-sharing pact initialed cohort I 0.78 166 OMITTED·EXPECTED
Fisheries quota ruling published cohort III 0.41 152 OMITTED·EXPECTED
Airport slots regulation amended cohort IV 0.19 147 OMITTED·EXPECTED
Pesticide reauthorization delayed cohort IV 11 UNDERPOWERED
Fig. 2. The instrument's target interface, on invented data. p₀ is the probability, under a null that preserves every outlet's weekly activity within its beat and every event's total reach, that the cohort would have covered the event anyway; the green tick marks the frozen 0.80 threshold. An omission is only surprising above it, a single event-cohort verdict is never more than secondary evidence, and too few comparable events yields no verdict at all. The aggregate line, not the named rows, is the study's primary estimand.

3The data

The matrix and every count on this page recompute from the downloadable extraction: incidence.csv, cohorts.json, summary.json, matrix-query.sql, and Fig. 2's invented values are marked as such in readout-specimen.json. The cohorts are provisional in the strict sense: membership is a TLD rule. A cohort like .com is mostly wires and international outlets, which is exactly why this page never calls cohorts markets. Earning that word takes a publisher-country audit of the full source panel against a frozen rule sheet, and that audit is scheduled as part of the preregistration, not retrofitted after it.

4Journal

2026-07 · why the interesting half is fake
A mockup that ran the null for real would be a pilot study wearing a costume, and its gaps would quietly shape the preregistration: thresholds picked after seeing which thresholds produce findings. So the real half of this page is confined to what the null treats as fixed, the incidence matrix and its margins, and the inferential half is confined to fiction. The scout pass and the go or no-go decision happen before any real verdict is computed.
2026-07 · what a verdict will never mean
OMITTED·SURPRISING is a statement about a frozen source panel, among events that panel observed, under a null preserving outlet activity and event reach. It is not a statement that an editor erred, that a story was objectively important, or that a country's media failed. Language stronger than that requires the audited source map and a source-availability record, and the eventual paper will say so in its limitations, not its footnotes.
2026-07 · what would kill it
Three things, all measurable. Degree-preserving randomization proving infeasible or degenerate at panel scale, in which case the null is theater. Event clustering precision too low, since under-merge and over-merge both corrupt the incidence matrix before any analysis touches it. And an unmapped residue of the source panel too large for cohorts to mean anything. Each gets a measured answer in the scout report before preregistration.

5Reported results

None yet. Under Rule 1, this experiment cannot appear in an edition until it has a bench entry. The first entry will be the preregistered omission-surprise run, whatever it finds, including nothing.