We took six data-rich events already published by the production pipeline and re-ran the writer step with two frontier models, Claude Fable 5 and GPT-5.6 Sol, under the production writer prompt with exactly two modifications: the chart-type whitelist (bar, line, pie, timeline) was replaced by an open vocabulary the writer names itself, and an explicit offline rule forbade web access. Inputs were identical to the original production runs: the same eight sorter-selected articles per event, same ordering, same clipping. We then audited every chart datapoint against the source articles and compared prose, headlines, and structure against the production output. Three results. The freedom is used: production drew all 12 of its charts from two forms, while the frontier writers produced 18 charts each across 12 distinct forms, most better matched to the story. The freedom is not abused: a datapoint audit found zero fabricated values in 152 frontier chart numbers. And in these six outputs the two models showed distinct editorial characters, neither dominant: Fable wrote richer journalism but broke the word-count ceiling in five of six events; Sol was the only writer with perfect format compliance but wrote flatter prose with a habit of inline over-attribution.
This abstract and its figures (two models, 152 audited datapoints) are the v1 study. A July 23 addendum adds a third writer, Kimi K3, on the same events and a full bench run; its numbers are reported there, not folded into the counts above.
Six events summarized by production between July 7 and July 10, 2026, chosen for numeric density and topic spread: a Greek two-pollster survey story, an OECD report on Italian employment and wages, the record June heatwave, the SK Hynix US listing, the Apollo counterbid for easyJet, and stage four of the Tour de France. All are multi-source (5 to 8 distinct outlets) and all received exactly eight articles after sorting.
Each model received the production system prompt (writer v2) rendered with the event's original summarization date and time, modified only in the CHARTS section (open type naming, flexible uniform data shape, a required one-sentence rendering note) plus an offline rule. Runs used production parameters (temperature 0.2) in plain JSON mode. The production baseline is the published English output of each event. "Production" below is the pipeline's current writer model; "Fable" is Claude Fable 5; "Sol" is GPT-5.6 Sol.
| Event | Production | Fable 5 | GPT-5.6 Sol |
|---|---|---|---|
| Greek polls | bar | grouped-bar, slope, bar ×2 | grouped-bar ×2, horizontal-bar ×2 |
| Italy OECD | bar | grouped-bar ×2, bar | grouped-bar ×2 |
| Hottest June | timeline, bar ×2 | bar ×3 | paired-dot-plot, horizontal-bar, comparison-bar |
| SK Hynix listing | bar ×3, timeline | bar, dumbbell, timeline | ranked-bar, comparison-bar, dumbbell |
| easyJet bid | timeline | dot-strip, timeline | bar ×2, timeline |
| Tour stage 4 | timeline, bar | dot-strip, timeline, bar | stat-strip, race-gap-board, timeline |
| Totals | 12 charts, 2 forms | 18 charts, 6 forms | 18 charts, 10 forms |
The interesting differences are editorial, not cosmetic:
Method: every numeric value in every chart was traced into the exact article clips the writer received (with European decimal and thousands formats normalized), and every non-match was resolved by hand against the source text.
Result: zero fabricated datapoints across all three writers (production 28 values, Fable 68, Sol 84). All initial misses resolved to one of: numbers written as words in the sources ("three minutes 50 seconds down", French "huitième jour"), correct and declared unit conversions (Fable's GC-gap bar states in its note that 473 seconds is 7:53), or formatting variants. The single most aggressive-looking prose claim, Fable crediting Apollo with a concurrent "move on Bayer", traces to a Bloomberg headline supplied in the input.
Two observations worth carrying to the bench. Cross-language numeracy was flawless in both models: values were extracted from Greek, German, French, Italian, Dutch, and Romanian text and normalized correctly. And both models converted units only when they declared the conversion in the chart note, which is exactly the behavior a grounding gate should require. These figures audit the open-vocabulary outputs of this experiment; the grounding column on the bench audits separate, closed-vocabulary reruns and is not the same measurement.
Erratum (July 10, evening). An external audit found a reference-baseline error in one Sol output: a temperature anomaly labeled against a "1991-2000" baseline where the other outputs and the underlying convention state 1991-2020. The value sits in a string field, which this audit's numeric matcher never inspected. The zero-fabrication result is therefore scoped to numeric chart values; string fields were not exhaustively audited, and the scope statements on this site have been corrected accordingly.
| Measure | Production | Fable 5 | GPT-5.6 Sol |
|---|---|---|---|
| Length inside the 450-600 word spec | 4/6 | 1/6 | 6/6 |
| Word range | 449-603 | 581-681 | 486-577 |
| Chart placeholders all valid | 6/6 | 6/6 | 6/6 |
| Renderer note on every chart | — | 18/18 | 18/18 |
| Em-dashes in prose | 0 | 0 | 0 |
| Banned-vocabulary hits | 4 | 0 | 0 |
| Quotes used, total | 13 | 14 | 7 |
Production tripped the brand banlist on four distinct terms, five occurrences ("major", "historic" twice, "highlights", "remarkable"); neither frontier model tripped it once. Fable's consistent overrun of the 600-word ceiling (five of six events, up to 681) is its one real compliance failure, and it is systematic rather than occasional. Sol under-uses quotation: several events with quotable material got a single quote where the prompt asks for one to three. Self-assigned quality scores also differ in character: Fable scores conservatively (5 to 7, matching the prompt's calibration anchors), Sol emits fractional scores and reached for an 8 on the heatwave event, which the anchors arguably permit but production and Fable both scored 7.
Reading all eighteen summaries side by side. One reader, six events: these are observations to replicate, not measured traits.
Fable writes like the stronger editor. It consistently recovered reporting the other two left in the sources: the Pulse poll's sample size and field dates (methodology production never mentioned), a Dutch corporate-finance professor's prediction that a low bid would "shake other bidders awake", translated verbatim per spec, and the UK insomnia survey in the Romanian coverage of the heatwave. Its section structure follows the story's logic rather than a template. The cost is discipline: it runs long, and the extra words are usually real facts, which makes the overrun harder to trim than padding would be.
Sol writes like a careful wire service. Sentences are precise, numbers are dense, lengths are exact, and nothing is speculative. Two habits hold it back editorially. It over-attributes inline ("according to ANSA", "in the Reuters report"), which is admirable grounding discipline expressed in the wrong register for a reader-facing summary. And it stays closer to enumeration than narrative; the heatwave summary reads as four accurate lists.
Production remains competent and is visibly the weakest of the three here. Its facts held up in this audit, but it missed the methodology details, the minor-source quotes, and the second poll's estimate figures, chose the generic chart every time, and was the only writer to violate the voice banlist.
Six events, one run per model, no statistical claims. The evaluator is a frontier model reading its own family's output; the grounding audit is mechanical and verifiable, but the prose judgments in section 5 are one reader's. Production output passed through post-processing that the experimental outputs did not. Costs were not measured on these runs and matter: the production writer is roughly an order of magnitude cheaper per article than either frontier model at list prices. The translator stage, which consumes charts downstream, was not exercised; open data shapes require translator schema changes before any of this ships.
Add the five recurring chart forms to the product renderer behind a flag, with the degrade map. Make the datapoint grounding audit a standing bench metric rather than a one-off. Re-run this six-event set with the length gate enforced and a quote-usage nudge to see whether each model's weakness is promptable or structural. Under Rule 1, nothing here enters an edition until the bench entry exists.
Kimi K3 (Moonshot AI) ran the same six events under the byte-identical prompts and the open chart vocabulary. Two run parameters deviate from the original protocol, both forced by how the model is served: its only endpoint does not accept a temperature, and it is a reasoning model whose thinking bills inside the completion budget. At the original 16,000-token cap, three of six runs spent the budget on reasoning before finishing the article; the cycling fixture needed a bounded reasoning setting after twice reasoning past even a doubled cap. The capped attempts are kept alongside the outputs in the experiment files.
The audited results match the frontier pattern: 18 charts across six forms (grouped-bar, slope, horizontal-bar, bar, timeline, and a fact-strip the renderer degrades to a board), and zero fabricated numeric chart values across 66 audited datapoints. One value is a disclosed derivation, a May poll standing reconstructed from the July figure and the sourced change, stated in the chart note; one is a definitional zero. No banned vocabulary. By the playground's word count, which strips block quotes and tags, all six summaries sit inside the 600-word band (474 to 582), so on these events Kimi does not repeat Fable's overruns; a raw count that includes Markdown headings and quote markers puts four nominally over, and the two counting methods are labeled as such wherever the numbers appear. Its outputs are in the playground as a fourth column. On the bench (run 2, all 13 fixtures, blind panel) it passes the grounding gate, clears 5/5 trap checks, and records the highest extraction recall observed so far, 83% — but under the open vocabulary, not run 1's closed prompt, so it is not a like-for-like result against the frontier writers' 78%. Judged form-fit 3.9, on conservative chart forms.
@techreport{pollarlabs006,
author = {Pollar Research},
title = {The writer with an open chart vocabulary},
year = {2026},
url = {https://labs.pollar.news/experiments/writer-charts/paper},
note = {Pollar Labs Experiment 006, v1.1 (July 23 addendum)}
}