31 · Creative and media · Compliance and trust · live
Editorial-control ledger
Eval results
Scored on a held-out or test splitRun 25 Sep 2026Eval write-up (decosa-api, access required)
- Edit metrics exact against scripted diffs80 / 80syntheticn = 80First run 59 / 80; a harness bug and a line-wrap bug were fixed before the second run.
- Rubber-stamped vs heavily edited separated (AUC, words changed)1.00syntheticn = 4332 rubber-stamped vs 11 heavy edits; the heavy edits are the builder's (3) and the model's (8), not newsroom editors'.
- Wording-only paraphrases with no claim change flagged (held out)24 / 24held outn = 24Generated after the last prompt change and never used to tune it.
- Claim-change flag on planted edits: recall / specificity24/24, 40/40syntheticn = 64The style-only row was seen during prompt tuning, so it is optimistic.
- IteraTeR test, meaning-changed edits flagged (recall)31 of 35 (0.89)test splitn = 35Precision against the intent label 0.49 (32 of 65 other edits flagged); the label is a proxy, not 'a claim changed'.
- Tampered ledgers caught120 / 120syntheticn = 12015 kinds of change on 8 ledgers; all 8 genuine ledgers verify.
Dataset
8 fictional newsroom source packs with a Qwen3.8-27B draft each, scripted and planted edits on them, plus the IteraTeR human revisions (Apache-2.0; dev split for prompt work, test split for the numbers).
Caveats
- No real newsroom editors on real copy: the heavy edits are the builder's and the model's.
- The style-only claim-diff cases were seen while tuning the prompt; only the 24 paraphrases are a fair wording-only number.
- Claim-diff precision was not measured on a set labelled for claim changes; IteraTeR intent labels are a proxy.
- A careful editor who changes nothing looks the same as a rubber stamp; the ledger does not judge review quality.
- Someone holding the server's signing key who also forges a fresh gateway receipt is not caught by the record alone.
- Pieces longer than about 24,000 characters and quiet-GPU latency are not measured.
Nightly smoke check
Loading the nightly status…
- Result
- partial
- Run
- 25 Sep 2026
- Latency, this run
- n/a
- p50 over passed runs
- 59 s
- Receipts
- 2
- Model calls
- n/a
- Tokens
- n/a
- Cost per run
- $0.002
Self-host verification
Verified on 25 Sep 2026: Fresh clone of the branch into a clean directory on our server, image built from docker/api/Dockerfile, api started with compose (named volume, python healthcheck), then the assemble prompt's smoke steps 1-8 and the CMS webhook with a minted dk_ key; torn down afterwards.
Verified 25 Sep 2026: image builds, the service starts healthy, and the sample passes end to end against local model servers equivalent to the documented ones (the already-running Qwen3.8-27B vLLM on 127.0.0.1:8114 instead of the compose llm service); model-server startup itself not re-verified. Receipts were attested (signed by the box's key). The C2PA credential answered 503 as documented: the default image has no c2pa-python and no certificate. Host networking and port 8437 were used because other services held the default ports.
Rehearsal bundle: editorial-ledger.zip (4 KB, 9 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.
Known limits
- Hosted check ran on the pre-release server (decosa-api the pre-release branch on our server, gateway route, signed receipts): console flow, public ledger page, text check, tamper buttons, Watch replay, 390 px layout, and the Build tab's Python example as written. The production API runs it once the branch is merged and deployed.
- The claim check is a model's reading: on IteraTeR it caught 31 of 35 meaning-changed edits and also fired on many 'clarity' edits (most of which did change a fact).
- The named editor's identity is what the caller sends; only an optional Ed25519 editor signature binds it to a key.
- The C2PA credential uses a development certificate (untrusted issuer) and needs the provenance extra in the image.
- Hosted retention is fixed at 7 days for open pieces and 30 days for published ledgers.
- Under heavy shared load a whole piece took up to a minute.
Models and licences
Standard tier. Licence posture: permissive (Apache, MIT or BSD).
- Ledger: edit metrics, sentence alignment, hash chain, sign-off rules, sealing and verification (no model; CPU)decosa-api editorial module (decosa_api/verticals/editorial)AGPL-3.0-or-later
- Writes the AI first draft from the sources, and lists the claims an edit added or removedQwen3.8-27B (NVFP4)Apache-2.0
- Optional C2PA content credential on a DOCX copy of the published textc2pa-python 0.37 (native c2pa-rs)MIT OR Apache-2.0
All quality evidence
Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.
Lite · ledger only, on CPU (self-host) (3)
- Edit metrics against scripted diffs of real drafts: 80 / 80 exactdocs/evals/editorial-ledger.md
- Rubber-stamped vs heavily edited, words changed: AUC 1.00; rubber-stamped at most 0.68%, heavy at least 51.1%docs/evals/editorial-ledger.md (32 rubber-stamped, 11 heavy)
- Tampered ledgers caught: 120 / 120 (15 kinds of change, 8 ledgers); 8 / 8 genuine verifieddocs/evals/editorial-ledger.md
Standard · Qwen3.8-27B drafts and checks claims (hosted demo) (4)
- Claim check on planted edits (number changed, fact deleted or added, paragraphs moved, style-only, 24 held-out paraphrases): 64 / 64; recall 24/24, no false alarm in 40docs/evals/editorial-ledger.md
- Claim check on IteraTeR human sentence edits (test): 31 / 35 meaning-changed caught; fired on 32 / 65 others (precision 0.49 against the intent label, which undercounts real fact changes)docs/evals/editorial-ledger.md
- Drafts with a gateway-signed receipt covering the stored text: 8 / 8 in the eval; 2 / 2 model calls per run in the end-to-end runsdocs/evals/editorial-ledger.md; scripts/smoke/editorial-ledger.py
- Edit metrics, separation and tamper detection: as the lite tier (same code)docs/evals/editorial-ledger.md