33 · Legal · live
Privileged drafting editor
Eval results
Scored on a held-out or test splitRun 25 Sep 2026Eval write-up (decosa-api, access required)
- Clause detection vs CUAD labels, v1 mapping (like for like): precision / recall0.77 / 0.79held outn = 3030 CUAD test contracts, 8 rules, run once. v1 was 0.78 / 0.82: no gain like for like.
- Clause detection vs CUAD labels, v2 mapping: precision / recall0.82 / 0.85held outn = 30Part of the gain is the rule redefinition (new CONSEQ rule matches CUAD 'Cap On Liability'), not better reading.
- Planted deviations flagged: precision / recall0.98 / 1.00held outn = 10252 of 52 deviations flagged, 1 false flag; 17 contracts could be planted.
- Exact position of a planted clause (5 classes)0.98held outn = 102100 / 102
- Inserted text not traceable to a precedent0 in 136 tracked changesheld outn = 136
- LibreOffice Accept All matches our proposal46 / 47held outn = 47Mismatch was a deleted last paragraph; writer bug fixed after the run, test set not re-run.
- Model cost per contract (median, list price)US$0.0089held out16 calls, 22.9k prompt + 1.5k output tokens
Dataset
CUAD v1 (510 SEC EDGAR contracts, expert clause labels, CC BY 4.0): 40 train-split dev contracts for tuning, 30 test-split contracts run once after the configuration was frozen. Playbook and precedents are synthetic; planted deviations from a clause-template bank.
Caveats
- The playbook and precedent library are synthetic, written for this demo; only clause presence and location come from CUAD.
- Per-rule samples are small (3 to 6 labelled contracts per rule): warranty duration found in 0 of 3 and insurance in 2 of 5.
- Only 17 of 30 test contracts had a buyer side, so only those could be planted.
- Not measured: real firm playbooks, Microsoft Word itself, agreement with a lawyer's own redline, latency on a dedicated card.
- One writer bug found by the test run was fixed afterwards; the test set was not re-run.
Nightly smoke check
Loading the nightly status…
- Result
- pass
- Run
- 25 Sep 2026
- Latency, this run
- n/a
- p50 over passed runs
- 28 s
- Receipts
- 20
- Model calls
- n/a
- Tokens
- n/a
- Cost per run
- $0.007
Self-host verification
Verified on 25 Sep 2026: Fresh clone of the branch into a clean directory, docker build of docker/api/Dockerfile, the assemble prompt's api service with a named volume, pointed at the already-running local vLLM (Qwen3.8-27B NVFP4 on 127.0.0.1:8114) through host networking; then torn down.
Verified on 2026-09-25: the image builds, the service starts healthy, the Northwind sample passes end to end (9.7 s, 20 attested calls, 7 tracked changes, 0 untraced insertions, the file's checks pass), the signed record verifies and fails when one finding is changed, and the .docx export downloads. The model server's own startup was not re-verified (no new GPU load).
Rehearsal bundle: legal-drafting-editor.zip (11 KB, 12 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.
Known limits
- The playbook and precedent library in the demo are synthetic; a firm's own need to be loaded (as JSON) and were not tested.
- Paragraphs with fields, hyperlinks, drawings or earlier tracked changes get a comment, not an in-place edit; inserted clauses are not numbered.
- Checked against the OOXML rules and LibreOffice; not yet opened in Microsoft Word by us.
- Position errors: 2% of planted clauses got the wrong one of five positions (v2 held-out run) and no planted deviation was missed; like-for-like clause detection against CUAD is 0.77 / 0.79, and warranty-duration and insurance clauses are the ones most often missed.
- The client-side playbook needs a buyer side: contracts between equals (joint ventures, cooperation agreements) get findings but party-specific precedents stay in comments.
Models and licences
Standard tier. Licence posture: permissive (Apache, MIT or BSD).
- Editor: DOCX reading and the tracked-change writer, candidate search, traceability check, file checks, signed record and ledger (no model; CPU)decosa-api drafting module (decosa_api/verticals/drafting)AGPL-3.0-or-later
- Model: places each clause against the playbook, proposes the find-and-replace edits, judges the grounding of each changed sentence, re-reads the edited clauseQwen3.8-27B (NVFP4)Apache-2.0
All quality evidence
Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.
Lite · one 32 GB card, self-hosted (1)
- Findings and redlines on CUAD: not measured separately: the same weights and prompts as the standard tier; speed on a 5090 not measurednot measured yet
Standard · one 96 GB card (measured; hosted demo) (9)
- Clause detection vs CUAD expert labels, 8 units × 30 held-out test contracts: precision / recall: 0.82 / 0.85 under the v2 mapping (v1: 0.81 / 0.79, re-scored); like for like with v1's mapping 0.77 / 0.79 (v1: 0.78 / 0.82), so no gain theredocs/evals/legal-drafting-editor.md (v2). The v2 gain is the new CONSEQ rule, which finds the exclusions of indirect damages CUAD files under "Cap On Liability" (9 of 9); warranty duration 0 of 3 and insurance 2 of 5 labelled contracts found
- Planted deviations flagged (102 planted clauses, 17 CUAD test contracts): precision / recall: 0.98 / 1.00 (52 of 52 deviations flagged, 1 false flag); exact position (of 5) 0.98docs/evals/legal-drafting-editor.md; v1 was 0.89 / 0.95, exact 0.85, on 13 contracts, and 0.95 / 0.97, exact 0.92, once 6 items a harness bug had left out of the file are removed
- Contracts where the buyer side was named (needed for party-specific precedents): 17 of 30 (v1: 13); both names among CUAD's labelled parties in 17 of 17docs/evals/legal-drafting-editor.md
- Redlines that are valid, reject-all = input, accept-all = proposal: 47 of 47docs/evals/legal-drafting-editor.md, every output file of the v2 held-out run
- Redlines LibreOffice opens and whose Accept All / Reject All match: 47 of 47 open, Reject All 47 of 47, Accept All 46 of 47; 284 of 284 comments importeddocs/evals/legal-drafting-editor.md: the mismatch was a deleted last paragraph (LibreOffice and Word keep an empty one); fixed after the run and tested, not re-run on the test set
- Inserted text not traceable to a firm precedent: 0 in 136 tracked changesdocs/evals/legal-drafting-editor.md; enforced in code: an untraced chunk sends the edit back to the precedent's sentence
- Targeted edits the checks sent back to the precedent's own sentence: 26 of 77docs/evals/legal-drafting-editor.md (v2 held-out run)
- Model cost per contract, eleven rules (median): US$0.0089 at list price ($0.30 / $1.50 per million tokens), 16 calls (v1: US$0.0065, 14 calls)docs/evals/legal-drafting-editor.md, 47 runs
- Levers tried on 40 dev contracts and left off: Self-consistency (3 readings): 2.6x the cost, no detection gain. Topic gate: +0.05 precision, -0.06 recall. Embedding retrieval (bge-small, CPU): +0 to +1 point of labelled text shown. Log-probabilities: the gateway does not return them.docs/evals/legal-drafting-editor.md, lever study
Wanted · a larger second reader on your own hardware (1)
- CUAD precision and recall on the same 8 rules: not measured yet