POLLAR LABS · BENCH · METHODOLOGY v0.4

The writer bench

Pollar Research
first results July 2026 · experiment 006 paper · playground
STATUS: TWO RUNS · 13 fixtures · four writers · blind judge panel

Abstract

The bench is where claims made by demos get checked: fixed fixtures, identical inputs per writer, published methodology, results reported with their costs and their misses. It separates gates from rankings. A gate (chart grounding, voice compliance) is something every serious writer must pass outright; reporting 100% next to 100% ranks nothing. Rankings (extraction recall, key-fact coverage, trap behavior, judged form-fit) are built to differ between good writers, including on adversarial fixtures designed to make the easy answer the wrong one. Under Rule 1, no experiment enters an edition without an entry here.

1Results, run 1 (July 10, 2026)

MeasureProduction writerClaude Fable 5GPT-5.6 Sol
Gate: voice and formatscreened this run (banned-vocabulary and limit counters include false positives); adjudicated pass/fail reporting starts with run 2
Gate: chart groundingflaggeda — 32/37pass — 81/81passb — 60/61
Extraction recall (core datasets)50%78%78%
Key-fact coveragec97%97%97%
Trap checks passed (4 synthetic fixtures, 5 checks)5/55/55/5
Form-fit, judgedd4.23.93.9
Summary length range (spec 450–600)212–705 words469–678421–675
Writer cost, full run$0.16$6.53$2.28
Table 1. Run 1 (tags bench-v05-*, July 10, 2026): 13 fixtures, identical frozen inputs, production writer prompt for all three writers, judged by the blind panel. Judge cost $0.55 total. Notes a–d below.

a Five values corrupted: clock times encoded as decimal numbers (a 3:50 gap charted as the value 3.5).

b The single unmatched value is the race leader's 0:00 gap, true by definition; adjudicated benign.

c Does not yet discriminate: the v0 key-fact probes (five to seven any-of text needles per fixture) are too coarse and are being tightened.

d Mean of per-chart medians; provisional. Under the production prompt all writers use the closed four-form chart vocabulary, so this measures execution within it, not form choice; the open-vocabulary comparison is in the 006 paper.

Reading notes. The gate worked in both directions. Production's general-classification chart encoded clock times as decimal numbers (a 3:50 gap charted as the value 3.5); five values were flagged and adjudicated as real corruption. Both frontier writers converted the same times to seconds correctly and declared the conversion in the chart note. On the traps, all three writers behaved: the tempting outlier figure (24 billion against a corroborated 2.4) was attributed, never echoed as fact, and the overlapping counts were not summed. In an earlier shakedown run production did assert a summed total as fact inside a chart, the bench's one true trap catch so far; that shakedown also exposed a draft of our trap scoring that failed writers for correctly attributing discrepancies, fixed before these results. Extraction recall separates writers most. Fact coverage does not discriminate yet and is marked accordingly rather than dressed up. Run identities: the numbers above are fresh closed-vocabulary runs (tags bench-v05-*); the earlier shakedown is bench-v04-first-run; the 006 paper audits a third, open-vocabulary set of outputs. At current fixture counts the rankings resolve production against the frontier writers, not the frontier writers against each other; a larger gold corpus is the path to finer resolution. The 2x2 is now complete: run bench-2x2-prod-open put the production model under the open chart vocabulary and its form range went from two to five (grouped bars, slopes, a horizontal bar), with a rendering note on every chart, while the same model's clock-time corruption did not recur. Together with the frontier writers collapsing to the whitelist under the closed prompt, vocabulary demonstrably drives form range in both directions; the frontier models still reach further (ten to twelve forms against five).

1bResults, run 2: a third frontier writer (July 23, 2026)

MeasureKimi K3 (open vocabulary)
Gate: voice and format0 banned vocabulary; one six-word ultra-short headline (band 3–5) in 13
Gate: chart groundingpasse — 108/110
Extraction recall (core datasets)f83%
Key-fact coveragec98%
Trap checks passed (4 synthetic fixtures, 5 checks)5/5
Form-fit, judgedd3.9
Summary length range (target 450–600)459–602 words; two of 13 a word or two over
Writer cost, full run$3.79
Table 2. Run 2 (tag bench-kimi-k3, July 23, 2026): the same 13 fixtures and blind panel, one writer, Kimi K3 (Moonshot AI) under the open chart vocabulary. Not comparable row-for-row to Table 1, which runs the closed production vocabulary; the closest published condition is bench-2x2-prod-open. Two run parameters were forced by how the model is served (no temperature on the endpoint; reasoning bills inside the completion budget, which capped shorter first attempts at the 16k token limit). Note e below.

e Two unmatched values, both disclosed derivations from sourced operands (a poll standing reconstructed from a reported change; a goal minute derived as opener plus a stated interval), each declared in its chart note and adjudicated benign. No fabricated numeric chart values.

f Under the open chart vocabulary; run 1's 78% for the frontier writers is under the closed production vocabulary, so this is the highest recall observed, not a like-for-like win. The comparable open-vocabulary condition is bench-2x2-prod-open.

Kimi K3 passes the grounding gate cleanly and behaves correctly on all four traps, including attributing the disagreeing casualty figures rather than picking one. On voice and format it carries no banned vocabulary and intact chart placeholders, with one genuine minor miss: a six-word ultra-short (push) headline against the three-to-five word band. It records the highest extraction recall observed so far, 83% of core datasets, though under the open vocabulary rather than run 1's closed prompt (note f); its judged form-fit sits with the others at 3.9. Its chart forms stay conservative under the open vocabulary, mostly bars with occasional grouped bars, slopes, and one fact strip, closer to the production model's reach than to the widest frontier outputs. Two of the thirteen summaries run a word or two past the 600-word target (601 and 602 by the stripped count); on the six playground events it stays inside the band. Kimi is a Moonshot model and therefore a contestant; Kimi K2.6 leaves the judge substitute pool while it is benched. Full outputs and per-fixture metrics are in the playground and the run artifacts.

2Metrics

Gates. Chart grounding: every numeric chart value must trace to the article text the writer received, through European number formats, clock durations (7:53 as 473 seconds), spelled-out numbers in seven languages, and composite durations ("three minutes 50 seconds"). Unmatched values are adjudicated by hand, not auto-failed; the gate passes only when every unmatched value is adjudicated benign, and denominators always include all values. Voice and format compliance: the brand banlist, headline limits, placeholder integrity, the word band.

Rankings. Extraction recall: each fixture carries a gold inventory of the distinct numeric datasets its sources contain; the score is the share of core datasets the writer actually charted. Key-fact coverage: hand-authored reportable facts (sample sizes, deadlines, records, minor-source quotes) matched against the prose. Trap behavior: pass or fail against documented correct behavior on adversarial fixtures. Form-fit: a blind three-judge panel scores every chart 1-5 on whether its visual form fits the data, with per-chart medians; judges see charts and a source-number digest, never the writer's identity.

3Fixtures

Thirteen, in three families. Six data-rich events from Experiment 006 with gold sets (frozen sorter selections reproduce the original production inputs). Three production snapshots without gold sets yet. Four synthetic trap fixtures carrying five scored checks (the disagreeing-sources fixture scores both figures), invented sources and events built around documented temptations: disagreeing casualty figures (both must appear, attributed), a magnitude clash between corroborated and sloppy sources, overlapping partial counts that must not be summed, and a sourced figure that contradicts model priors. Trap content is never real news and is labeled as such in the fixture files.

4Judges

The panel is drawn from companies with no writer in the contest: Qwen 3.7-Max (Alibaba), Gemini 3.5 Flash (Google), Grok 4.5 (xAI), the strongest available non-contestant models as of July 2026. Judging is blind and rubric-bound; per-judge verdicts and comments are part of the recorded result. The panel costs about 1.4 cents per judged output, so it runs on every bench pass.

5Reporting rules

Saturated metrics are reported as gates, never as ranking bars. Trap fixtures are labeled synthetic wherever scores appear. Judge identity and rubric version publish with the scores; per-judge verdicts, gold sets, and the adjudication log publish in the artifact bundle below, gaps stated where they exist. Run costs publish with the results. Results that embarrass an experiment publish anyway; that is what the bench is for.

6Artifacts

The full bundle behind the numbers above: README (inventory and known gaps), adjudication log, judge rubric, fixture provenance, cost ledger. Runs, each as raw outputs (.json) and full report (.md): v05-prod (json), v05-fable (json), v05-sol (json), 2x2-prod-open (json), kimi-k3 (json), v04 shakedown (json). Gold sets and the four synthetic trap fixtures (content invented, published in full) are in the same directory; real-event article text is withheld for copyright and listed by source and URL in the provenance file.