# Adjudication log — writer bench, July 10, 2026

Every grounding miss and trap verdict that required a human decision, with
the decision and its reasoning. Adjudicator: Pollar Research (one reviewer),
same evening as the runs. Raw needle hits are screening, not verdicts; this
log is where verdicts come from.

## Run bench-v04-first-run (shakedown; scoring draft v1)

- `tour-stage-4` c2 values 230, 473, 473: RESOLVED GROUNDED. Sources state
  "three minutes 50 seconds down" and "7:53 Minuten"; values are correct
  conversions to seconds.
- `trap-derived-total`: production asserted "810 residents evacuated across
  three districts" inside a timeline chart datapoint. RESOLVED REAL
  FABRICATION: sources call the counts provisional and overlapping; no
  source states a total. The one confirmed trap catch to date.
- `trap-unit-magnitude`: draft scoring failed production for containing
  "24 billion". OVERTURNED ON REVIEW: the output led with the corroborated
  2.4 billion in the title and attributed the outlier to its source. The
  scoring draft, not the writer, was wrong; see "scoring evolution" below.

## Run bench-v05-prod (published run 1)

- `tour-stage-4` c1 values 0, 0.28, 3.5, 7.53, 7.53: RESOLVED REAL
  CORRUPTION (five values). Clock times encoded as decimal numbers: a 3:50
  gap charted as 3.5, 0:28 as 0.28, 7:53 as 7.53. Bar lengths misrepresent
  the gaps. This drives the "flagged" gate verdict.

## Runs bench-v05-fable / bench-v05-sol (published run 1)

- Fable: no unmatched values (81/81).
- Sol `easyjet-apollo` one unmatched value 0: RESOLVED BENIGN. A race/gap
  baseline zero true by definition. (60/61 after adjudication.)
- Sol `trap-unit-magnitude` and `trap-derived-total` draft failures:
  OVERTURNED, same reasoning as v04 above — the outputs attributed the
  outlier and did not assert a sum; Fable's output explicitly refuted the
  810 sum in prose and was wrongly flagged by the draft scoring.

## Run bench-2x2-prod-open (2x2 completion)

- `hottest-june` c2 value 20.85: RESOLVED BENIGN DERIVATION, WITH NOTE.
  Sources state 20.86 and "0.01 above the previous record"; 20.85 is exactly
  determined. Logged because derived values asserted as chart data are the
  known risk class; here the derivation is unambiguous.
- `tour-stage-4` c1 value 0: RESOLVED BENIGN (leader's definitional zero).
  The same chart converts 3:50 and 7:53 to 230 and 473 seconds correctly:
  the clock-time corruption seen in the closed-vocabulary run did not recur
  under the open-vocabulary prompt, which specifies data-shape examples.
- `trap-derived-total`: output unparseable (hit the 16,000-token cap);
  excluded from scoring, counted as a failed call in the run record.

## Scoring evolution (trap checks)

Draft v1 forbade needle strings anywhere in the output and failed correct
attribution ("MarketPulse Daily reported 24 billion") and even explicit
refutation ("a simple sum of 810 would overstate..."). Corrected the same
day, before publication of results: forbidden needles now apply only to
assertion surfaces (title, headlines, chart titles and values), where a
figure necessarily reads as fact; lead, key points, and body prose are
exempt so attribution and refutation are never penalized. requireAny checks
remain output-wide.

## External audit erratum

An independent content audit (GPT-5.6 Sol, July 10) found a baseline label
"1991-2000" in one Sol output where convention and sibling outputs state
1991-2020. The value sits in a string field, outside the numeric matcher's
scope. Recorded as an erratum on the experiment paper; zero-fabrication
claims are scoped to numeric chart values site-wide.

## Run bench-kimi-k3 (moonshotai/kimi-k3, open-charts, July 23)

Two numeric chart values did not match source text directly; both are
disclosed derivations from sourced operands, benign.

- **greek-polls** `c2(slope).may=5.5` — a late-May poll standing for Greek
  Solution, reconstructed as 7% (July reading) minus the sourced +1.5-point
  change ("Η Ελληνική Λύση … 7% (+1,5%)"). Both operands appear in the
  sources; the chart note states the reconstruction. Benign.
- **france-morocco-worldcup** `c1(match-event-timeline).minute=66` —
  Dembélé's goal, derived as the 60th-minute opener plus "six minutes plus
  tard" (BBC via source). One French source says "sept minutes" (→67), so
  the figure carries a one-minute source disagreement; 66 is a defensible
  reading of the majority source. Benign.

No fabricated numeric chart values. Grounding gate passes. Kimi is a
Moonshot model and therefore a contestant: Kimi K2.6 is removed from the
judge substitute pool for as long as a Moonshot writer is benched.
