Eval: editorial-control ledger (vertical 31)
Run 25 Sep 2026 on our server. Model: Qwen3.8-27B through the model gateway (every call has a gateway-signed receipt),
temperature 0, thinking off, while the GPU was shared with other evaluation jobs. Script: scripts/editorial_eval.py;
raw results: docs/evals/editorial-ledger-results.json; cached model outputs: docs/evals/editorial-ledger/.
Data:
- 8 fictional newsroom source packs (the two console samples plus six in
editorial-ledger/sources.json, all invented names and figures). The model wrote one AI draft for each (8 of 8 with signed receipts; mean 556 generated tokens, p50 8.6 s). - IteraTeR (Du et al. 2022, Apache-2.0,
wanyu/IteraTeR_human_sentand_human_doc): human revisions of arXiv, Wikipedia and news sentences and documents with edit-intent labels. Dev split for prompt work, test split for the numbers below.
1. Edit metrics against known diffs: 80 of 80 exact
Scripted edits on the 8 real drafts with exact expected counts: 1, 3, 10 and 25 word substitutions; 1 and 2 sentence deletions; 1 and 2 sentence insertions; one sentence edited with 2 substitutions; a reflow that only changes spacing and line breaks. Checked: deleted and inserted words, word-level edit distance, and sentence counts (removed, added, edited, and the edited sentence's own distance).
| Edit | Exact |
|---|---|
| substitutions (1, 3, 10, 25 words) | 32 / 32 |
| sentence deletions (1, 2) | 16 / 16 |
| sentence insertions (1, 2) | 16 / 16 |
| one sentence edited, distance 2 | 8 / 8 |
| spacing and line-break reflow (all zero) | 8 / 8 |
First run: 59 of 80. Two causes, both fixed before the second run: (a) the harness substituted capitalised
sentence-initial words with lowercase tokens, which genuinely merges two sentences for the splitter (a test bug: it now
substitutes only lowercase mid-sentence words); (b) a single line break inside a paragraph counted as a sentence end, so
a reflow looked like 3 added sentences (a real bug: metrics._unwrap now treats it as wrapping; blank lines and list
bullets still split). Word-level counts were exact in both runs.
2. Rubber-stamped vs heavily edited: fully separated
- Rubber-stamped (32): the draft unchanged, reflowed, a punctuation fix, or one word changed.
- Heavily edited (11): 3 hand edits of real drafts by the builder (not a newsroom editor) and 8 rewrites by the model prompted as a demanding editor (new lede, cut a third, restructure). Model-simulated edits are a stand-in for people.
- For reference, 51 real human document revisions from IteraTeR (test).
| Metric | Rubber-stamped (min / median / max) | Heavy | IteraTeR human docs (median) | AUC heavy vs rubber |
|---|---|---|---|---|
| words changed | 0 / 0 / 0.68% | 51.1 / 70.7 / 77.8% | 10.1% | 1.00 |
| published words not from the draft | 0 / 0 / 0.68% | 28.3 / 50.0 / 63.0% | 10.9% | 1.00 |
| sentence edit distance | 0 / 0 / 0.009 | 0.80 / 0.94 / 1.00 | 0.20 | 1.00 |
Depth bands: rubber-stamped 19 "none" and 13 "light"; heavy 11 of 11 "heavy"; IteraTeR docs 11 light, 23 moderate, 17 heavy. So the bands separate the extremes easily; real edits sit in between, where the numbers are a description, not a verdict. The obvious limit: a careful editor who reads everything and changes nothing looks the same as a rubber stamp. That is why the ledger also records time and a named sign-off, and why it never claims to judge review quality.
3. Claims added and removed (the model's claim diff)
Planted edits on the 8 drafts, run through the real pipeline (alignment, then one receipted call with both articles as context). A case is right when "any removed" and "any added" both match the expectation.
| Planted edit | Expect | Right |
|---|---|---|
| a number changed | removed and added | 8 / 8 |
| a fact sentence deleted | removed only | 8 / 8 |
| a fact sentence added | added only | 8 / 8 |
| two paragraphs swapped (facts only moved) | none | 8 / 8 |
| style words swapped (said/stated, will/is set to, However/But) | none | 8 / 8 |
| one sentence paraphrased by the model, facts kept (held out) | none | 24 / 24 |
Claim-change flag over all 64: recall 24/24, specificity 40/40. 64 calls, 64 signed receipts, no unparsed answers; p50 12.3 s, p90 17.9 s per call on the shared gateway; about 1,790 prompt and 57 generated tokens per call.
Honesty on tuning: the first run used an earlier prompt, which flagged 6 of the 8 style-only cases (it listed the
reworded sentence as one claim removed and one added). I then added a same_facts field and a reworded-claim filter,
and the prompt names said/stated and will/is set to as examples, so the style-only row above is optimistic: it was seen.
The prompt was otherwise checked only on IteraTeR dev (80 edits). The 24 paraphrases were generated after that change
and never used to tune it; they are the fair number for wording-only edits.
IteraTeR test (35 "meaning-changed" edits and a seeded sample of 65 others): the flag fired on 31 of 35 meaning-changed (recall 0.89) and on 32 of 65 others (precision 0.49 against the intent label). By label: clarity 29/41, fluency 1/14, coherence 1/6, style 1/3. The intent label is not "a claim changed": in a manual audit of 12 flagged "clarity" edits by the builder, 10 did add or remove a checkable statement (a dropped "(23 months)", "Shannon entropy" removed from a list, different models compared), 1 was unclear and 1 was wording only. So precision against true claim changes is well above 0.49, but it was not measured on a labelled set; treat the claim list as a pointer for the editor.
4. Tamper detection: 120 of 120
Ledgers built from the 8 real drafts with their gateway receipts, two revisions each (a light one and the heavy one), a real claim check and a sign-off. 15 kinds of change, each on all 8:
- in place (the signature is left alone): draft text edited, revision text edited, sign-off name changed, label flipped, a revision deleted, entries reordered, metrics inflated, claims emptied, the draft's gateway receipt pointed at other text, the signed summary's final hash changed, the signature forged;
- rewritten and re-signed with another key (every hash, link, root and signature recomputed): metrics inflated, draft replaced, sign-off moved to an earlier version, a revision dropped.
All 120 fail verification under the ledger's own rules, without relying on key pinning: re-signed records are caught because the metrics recompute from the texts, the draft's text must match its gateway receipt, and each revision must build on the previous version's hash. All 8 genuine ledgers verify. Limit: someone who holds the server's signing key and rewrites the draft together with a fresh gateway receipt is not caught by this record alone; that needs the gateway's own receipt log or an outside timestamp.
What is not measured
- Real newsroom editors on real copy (the heavy edits are the builder's and the model's).
- Claim-diff precision on a set labelled for claim changes; the IteraTeR intent labels are a proxy.
- Pieces longer than about 24,000 characters in both versions, where the claim check drops the full-article context.
- Latency when the GPU is quiet: all numbers above are under shared load.