EXPERIMENT 006 · PAPER · v1.1

The writer with an open chart vocabulary

Pollar Research
July 2026 · figures live in the playground
STATUS: MOCKUP · v1 two-writer study · July 23 addendum adds a third · bench: runs 1–2

Abstract

We took six data-rich events already published by the production pipeline and re-ran the writer step with two frontier models, Claude Fable 5 and GPT-5.6 Sol, under the production writer prompt with exactly two modifications: the chart-type whitelist (bar, line, pie, timeline) was replaced by an open vocabulary the writer names itself, and an explicit offline rule forbade web access. Inputs were identical to the original production runs: the same eight sorter-selected articles per event, same ordering, same clipping. We then audited every chart datapoint against the source articles and compared prose, headlines, and structure against the production output. Three results. The freedom is used: production drew all 12 of its charts from two forms, while the frontier writers produced 18 charts each across 12 distinct forms, most better matched to the story. The freedom is not abused: a datapoint audit found zero fabricated values in 152 frontier chart numbers. And in these six outputs the two models showed distinct editorial characters, neither dominant: Fable wrote richer journalism but broke the word-count ceiling in five of six events; Sol was the only writer with perfect format compliance but wrote flatter prose with a habit of inline over-attribution.

This abstract and its figures (two models, 152 audited datapoints) are the v1 study. A July 23 addendum adds a third writer, Kimi K3, on the same events and a full bench run; its numbers are reported there, not folded into the counts above.

1Setup

Six events summarized by production between July 7 and July 10, 2026, chosen for numeric density and topic spread: a Greek two-pollster survey story, an OECD report on Italian employment and wages, the record June heatwave, the SK Hynix US listing, the Apollo counterbid for easyJet, and stage four of the Tour de France. All are multi-source (5 to 8 distinct outlets) and all received exactly eight articles after sorting.

Each model received the production system prompt (writer v2) rendered with the event's original summarization date and time, modified only in the CHARTS section (open type naming, flexible uniform data shape, a required one-sentence rendering note) plus an offline rule. Runs used production parameters (temperature 0.2) in plain JSON mode. The production baseline is the published English output of each event. "Production" below is the pipeline's current writer model; "Fable" is Claude Fable 5; "Sol" is GPT-5.6 Sol.

2Chart forms

EventProductionFable 5GPT-5.6 Sol
Greek pollsbargrouped-bar, slope, bar ×2grouped-bar ×2, horizontal-bar ×2
Italy OECDbargrouped-bar ×2, bargrouped-bar ×2
Hottest Junetimeline, bar ×2bar ×3paired-dot-plot, horizontal-bar, comparison-bar
SK Hynix listingbar ×3, timelinebar, dumbbell, timelineranked-bar, comparison-bar, dumbbell
easyJet bidtimelinedot-strip, timelinebar ×2, timeline
Tour stage 4timeline, bardot-strip, timeline, barstat-strip, race-gap-board, timeline
Totals12 charts, 2 forms18 charts, 6 forms18 charts, 10 forms
Table 1. Chart output per event and writer. Every chart renders live in the playground.

The interesting differences are editorial, not cosmetic:

  • The story got charted, not just the numbers. The Greek event's headline is that the gap narrowed. Production charted a single-poll snapshot. Fable charted the movement itself: a slope chart from the late-May measurement to July, plus a grouped-bar putting both pollsters side by side, which is the honest way to report a poll story. Sol charted intention and estimate as two pollster-paired panels.
  • Comparisons found their form. Both frontier models independently chose a dumbbell for the Micron vs SK Hynix valuation gap (strictly a two-point dot comparison; each company carries one value). Fable used a dot-strip to place easyJet's share price under both takeover offers on one line, replacing three sentences of arithmetic.
  • Context became a figure. Fable produced a dot-strip of every Norwegian to wear the Tour's yellow jersey, every point traceable to the articles. Production had no form available for that idea.
  • Uncertainty entered the schema. Sol encoded wildfire figures as a lower bound plus a "more than" qualifier and kept each temperature figure's stated baseline beside it. This is data honesty the current schema cannot express.
  • Sol also invented bespoke forms (race-gap-board, stat-strip) that a renderer must interpret; its notes carry enough instruction to do so.

3Grounding audit

Method: every numeric value in every chart was traced into the exact article clips the writer received (with European decimal and thousands formats normalized), and every non-match was resolved by hand against the source text.

Result: zero fabricated datapoints across all three writers (production 28 values, Fable 68, Sol 84). All initial misses resolved to one of: numbers written as words in the sources ("three minutes 50 seconds down", French "huitième jour"), correct and declared unit conversions (Fable's GC-gap bar states in its note that 473 seconds is 7:53), or formatting variants. The single most aggressive-looking prose claim, Fable crediting Apollo with a concurrent "move on Bayer", traces to a Bloomberg headline supplied in the input.

Two observations worth carrying to the bench. Cross-language numeracy was flawless in both models: values were extracted from Greek, German, French, Italian, Dutch, and Romanian text and normalized correctly. And both models converted units only when they declared the conversion in the chart note, which is exactly the behavior a grounding gate should require. These figures audit the open-vocabulary outputs of this experiment; the grounding column on the bench audits separate, closed-vocabulary reruns and is not the same measurement.

Erratum (July 10, evening). An external audit found a reference-baseline error in one Sol output: a temperature anomaly labeled against a "1991-2000" baseline where the other outputs and the underlying convention state 1991-2020. The value sits in a string field, which this audit's numeric matcher never inspected. The zero-fabrication result is therefore scoped to numeric chart values; string fields were not exhaustively audited, and the scope statements on this site have been corrected accordingly.

4Format and voice compliance

MeasureProductionFable 5GPT-5.6 Sol
Length inside the 450-600 word spec4/61/66/6
Word range449-603581-681486-577
Chart placeholders all valid6/66/66/6
Renderer note on every chart18/1818/18
Em-dashes in prose000
Banned-vocabulary hits400
Quotes used, total13147
Table 2. Compliance with the writer prompt's format and voice rules.

Production tripped the brand banlist on four distinct terms, five occurrences ("major", "historic" twice, "highlights", "remarkable"); neither frontier model tripped it once. Fable's consistent overrun of the 600-word ceiling (five of six events, up to 681) is its one real compliance failure, and it is systematic rather than occasional. Sol under-uses quotation: several events with quotable material got a single quote where the prompt asks for one to three. Self-assigned quality scores also differ in character: Fable scores conservatively (5 to 7, matching the prompt's calibration anchors), Sol emits fractional scores and reached for an 8 on the heatwave event, which the anchors arguably permit but production and Fable both scored 7.

5Prose

Reading all eighteen summaries side by side. One reader, six events: these are observations to replicate, not measured traits.

Fable writes like the stronger editor. It consistently recovered reporting the other two left in the sources: the Pulse poll's sample size and field dates (methodology production never mentioned), a Dutch corporate-finance professor's prediction that a low bid would "shake other bidders awake", translated verbatim per spec, and the UK insomnia survey in the Romanian coverage of the heatwave. Its section structure follows the story's logic rather than a template. The cost is discipline: it runs long, and the extra words are usually real facts, which makes the overrun harder to trim than padding would be.

Sol writes like a careful wire service. Sentences are precise, numbers are dense, lengths are exact, and nothing is speculative. Two habits hold it back editorially. It over-attributes inline ("according to ANSA", "in the Reuters report"), which is admirable grounding discipline expressed in the wrong register for a reader-facing summary. And it stays closer to enumeration than narrative; the heatwave summary reads as four accurate lists.

Production remains competent and is visibly the weakest of the three here. Its facts held up in this audit, but it missed the methodology details, the minor-source quotes, and the second poll's estimate figures, chose the generic chart every time, and was the only writer to violate the voice banlist.

6Implications

  • The chart whitelist, not the model, is the binding constraint on form range; this is now measured across the full grid, not inferred from the confounded original design. Both frontier writers under the closed production prompt collapsed to the whitelist (bar, timeline, one pie, one line across 65 charts); the production model under the open vocabulary (bench run 2x2-prod-open) went from two forms to five, with a rendering note on every chart. Vocabulary moves form range in both directions for all three models. Model still governs reach: five forms for production against ten to twelve for the frontier writers. And opening the vocabulary cost zero fabricated numeric values in any writer.
  • A small renderer vocabulary covers almost everything produced. Grouped-bar, slope, dumbbell, dot-strip, and horizontal-bar, added to the existing four, cover 30 of the 36 frontier charts by exact name and 34 once evident families fold (comparison-bar and ranked-bar are bars; paired-dot-plot is a dot form); the remaining bespoke forms (stat-strip, race-gap-board) degrade to typographic boards using the data alone.
  • The rendering note earns its place. Every one of the 36 charts carried a usable instruction; several (unit conversions, lower bounds, baselines) contain information the renderer must have to be honest.
  • Sol's qualifier encoding is worth adopting regardless of model choice: a lower bound plus qualifier is how "more than 35,000 hectares" should be charted.
  • Model choice is a genuine editorial decision. Fable for depth and voice with a hard length gate bolted on; Sol for compliance and precision with an attribution-style rule and a quote-usage nudge. This is what the bench exists to arbitrate.

7Limitations

Six events, one run per model, no statistical claims. The evaluator is a frontier model reading its own family's output; the grounding audit is mechanical and verifiable, but the prose judgments in section 5 are one reader's. Production output passed through post-processing that the experimental outputs did not. Costs were not measured on these runs and matter: the production writer is roughly an order of magnitude cheaper per article than either frontier model at list prices. The translator stage, which consumes charts downstream, was not exercised; open data shapes require translator schema changes before any of this ships.

8Next

Add the five recurring chart forms to the product renderer behind a flag, with the degrade map. Make the datapoint grounding audit a standing bench metric rather than a one-off. Re-run this six-event set with the length gate enforced and a quote-usage nudge to see whether each model's weakness is promptable or structural. Under Rule 1, nothing here enters an edition until the bench entry exists.

AAddendum: a third frontier writer (23 July 2026)

Kimi K3 (Moonshot AI) ran the same six events under the byte-identical prompts and the open chart vocabulary. Two run parameters deviate from the original protocol, both forced by how the model is served: its only endpoint does not accept a temperature, and it is a reasoning model whose thinking bills inside the completion budget. At the original 16,000-token cap, three of six runs spent the budget on reasoning before finishing the article; the cycling fixture needed a bounded reasoning setting after twice reasoning past even a doubled cap. The capped attempts are kept alongside the outputs in the experiment files.

The audited results match the frontier pattern: 18 charts across six forms (grouped-bar, slope, horizontal-bar, bar, timeline, and a fact-strip the renderer degrades to a board), and zero fabricated numeric chart values across 66 audited datapoints. One value is a disclosed derivation, a May poll standing reconstructed from the July figure and the sourced change, stated in the chart note; one is a definitional zero. No banned vocabulary. By the playground's word count, which strips block quotes and tags, all six summaries sit inside the 600-word band (474 to 582), so on these events Kimi does not repeat Fable's overruns; a raw count that includes Markdown headings and quote markers puts four nominally over, and the two counting methods are labeled as such wherever the numbers appear. Its outputs are in the playground as a fourth column. On the bench (run 2, all 13 fixtures, blind panel) it passes the grounding gate, clears 5/5 trap checks, and records the highest extraction recall observed so far, 83% — but under the open vocabulary, not run 1's closed prompt, so it is not a like-for-like result against the frontier writers' 78%. Judged form-fit 3.9, on conservative chart forms.

9Cite

@techreport{pollarlabs006,
  author = {Pollar Research},
  title  = {The writer with an open chart vocabulary},
  year   = {2026},
  url    = {https://labs.pollar.news/experiments/writer-charts/paper},
  note   = {Pollar Labs Experiment 006, v1.1 (July 23 addendum)}
}