Skip to content
decosa

The full write-up behind the numbers on the tool’s page, as the team that built it wrote it: data, method, results and caveats. Internal names are removed; nothing else is edited.

Eval: editorial-control ledger (vertical 31)

Run 25 Sep 2026 on our server. Model: Qwen3.8-27B through the model gateway (every call has a gateway-signed receipt), temperature 0, thinking off, while the GPU was shared with other evaluation jobs. Script: scripts/editorial_eval.py; raw results: docs/evals/editorial-ledger-results.json; cached model outputs: docs/evals/editorial-ledger/.

Data:

  • 8 fictional newsroom source packs (the two console samples plus six in editorial-ledger/sources.json, all invented names and figures). The model wrote one AI draft for each (8 of 8 with signed receipts; mean 556 generated tokens, p50 8.6 s).
  • IteraTeR (Du et al. 2022, Apache-2.0, wanyu/IteraTeR_human_sent and _human_doc): human revisions of arXiv, Wikipedia and news sentences and documents with edit-intent labels. Dev split for prompt work, test split for the numbers below.

1. Edit metrics against known diffs: 80 of 80 exact

Scripted edits on the 8 real drafts with exact expected counts: 1, 3, 10 and 25 word substitutions; 1 and 2 sentence deletions; 1 and 2 sentence insertions; one sentence edited with 2 substitutions; a reflow that only changes spacing and line breaks. Checked: deleted and inserted words, word-level edit distance, and sentence counts (removed, added, edited, and the edited sentence's own distance).

Edit Exact
substitutions (1, 3, 10, 25 words) 32 / 32
sentence deletions (1, 2) 16 / 16
sentence insertions (1, 2) 16 / 16
one sentence edited, distance 2 8 / 8
spacing and line-break reflow (all zero) 8 / 8

First run: 59 of 80. Two causes, both fixed before the second run: (a) the harness substituted capitalised sentence-initial words with lowercase tokens, which genuinely merges two sentences for the splitter (a test bug: it now substitutes only lowercase mid-sentence words); (b) a single line break inside a paragraph counted as a sentence end, so a reflow looked like 3 added sentences (a real bug: metrics._unwrap now treats it as wrapping; blank lines and list bullets still split). Word-level counts were exact in both runs.

2. Rubber-stamped vs heavily edited: fully separated

  • Rubber-stamped (32): the draft unchanged, reflowed, a punctuation fix, or one word changed.
  • Heavily edited (11): 3 hand edits of real drafts by the builder (not a newsroom editor) and 8 rewrites by the model prompted as a demanding editor (new lede, cut a third, restructure). Model-simulated edits are a stand-in for people.
  • For reference, 51 real human document revisions from IteraTeR (test).
Metric Rubber-stamped (min / median / max) Heavy IteraTeR human docs (median) AUC heavy vs rubber
words changed 0 / 0 / 0.68% 51.1 / 70.7 / 77.8% 10.1% 1.00
published words not from the draft 0 / 0 / 0.68% 28.3 / 50.0 / 63.0% 10.9% 1.00
sentence edit distance 0 / 0 / 0.009 0.80 / 0.94 / 1.00 0.20 1.00

Depth bands: rubber-stamped 19 "none" and 13 "light"; heavy 11 of 11 "heavy"; IteraTeR docs 11 light, 23 moderate, 17 heavy. So the bands separate the extremes easily; real edits sit in between, where the numbers are a description, not a verdict. The obvious limit: a careful editor who reads everything and changes nothing looks the same as a rubber stamp. That is why the ledger also records time and a named sign-off, and why it never claims to judge review quality.

3. Claims added and removed (the model's claim diff)

Planted edits on the 8 drafts, run through the real pipeline (alignment, then one receipted call with both articles as context). A case is right when "any removed" and "any added" both match the expectation.

Planted edit Expect Right
a number changed removed and added 8 / 8
a fact sentence deleted removed only 8 / 8
a fact sentence added added only 8 / 8
two paragraphs swapped (facts only moved) none 8 / 8
style words swapped (said/stated, will/is set to, However/But) none 8 / 8
one sentence paraphrased by the model, facts kept (held out) none 24 / 24

Claim-change flag over all 64: recall 24/24, specificity 40/40. 64 calls, 64 signed receipts, no unparsed answers; p50 12.3 s, p90 17.9 s per call on the shared gateway; about 1,790 prompt and 57 generated tokens per call.

Honesty on tuning: the first run used an earlier prompt, which flagged 6 of the 8 style-only cases (it listed the reworded sentence as one claim removed and one added). I then added a same_facts field and a reworded-claim filter, and the prompt names said/stated and will/is set to as examples, so the style-only row above is optimistic: it was seen. The prompt was otherwise checked only on IteraTeR dev (80 edits). The 24 paraphrases were generated after that change and never used to tune it; they are the fair number for wording-only edits.

IteraTeR test (35 "meaning-changed" edits and a seeded sample of 65 others): the flag fired on 31 of 35 meaning-changed (recall 0.89) and on 32 of 65 others (precision 0.49 against the intent label). By label: clarity 29/41, fluency 1/14, coherence 1/6, style 1/3. The intent label is not "a claim changed": in a manual audit of 12 flagged "clarity" edits by the builder, 10 did add or remove a checkable statement (a dropped "(23 months)", "Shannon entropy" removed from a list, different models compared), 1 was unclear and 1 was wording only. So precision against true claim changes is well above 0.49, but it was not measured on a labelled set; treat the claim list as a pointer for the editor.

4. Tamper detection: 120 of 120

Ledgers built from the 8 real drafts with their gateway receipts, two revisions each (a light one and the heavy one), a real claim check and a sign-off. 15 kinds of change, each on all 8:

  • in place (the signature is left alone): draft text edited, revision text edited, sign-off name changed, label flipped, a revision deleted, entries reordered, metrics inflated, claims emptied, the draft's gateway receipt pointed at other text, the signed summary's final hash changed, the signature forged;
  • rewritten and re-signed with another key (every hash, link, root and signature recomputed): metrics inflated, draft replaced, sign-off moved to an earlier version, a revision dropped.

All 120 fail verification under the ledger's own rules, without relying on key pinning: re-signed records are caught because the metrics recompute from the texts, the draft's text must match its gateway receipt, and each revision must build on the previous version's hash. All 8 genuine ledgers verify. Limit: someone who holds the server's signing key and rewrites the draft together with a fresh gateway receipt is not caught by this record alone; that needs the gateway's own receipt log or an outside timestamp.

What is not measured

  • Real newsroom editors on real copy (the heavy edits are the builder's and the model's).
  • Claim-diff precision on a set labelled for claim changes; the IteraTeR intent labels are a proxy.
  • Pieces longer than about 24,000 characters in both versions, where the claim check drops the full-article context.
  • Latency when the GPU is quiet: all numbers above are under shared load.