Skip to content
decosa

33 · Legal · live

Privileged drafting editor

Open the toolJSON

Eval results

Scored on a held-out or test splitRun 25 Sep 2026Eval write-up (decosa-api, access required)

  • Clause detection vs CUAD labels, v1 mapping (like for like): precision / recall0.77 / 0.79held outn = 3030 CUAD test contracts, 8 rules, run once. v1 was 0.78 / 0.82: no gain like for like.
  • Clause detection vs CUAD labels, v2 mapping: precision / recall0.82 / 0.85held outn = 30Part of the gain is the rule redefinition (new CONSEQ rule matches CUAD 'Cap On Liability'), not better reading.
  • Planted deviations flagged: precision / recall0.98 / 1.00held outn = 10252 of 52 deviations flagged, 1 false flag; 17 contracts could be planted.
  • Exact position of a planted clause (5 classes)0.98held outn = 102100 / 102
  • Inserted text not traceable to a precedent0 in 136 tracked changesheld outn = 136
  • LibreOffice Accept All matches our proposal46 / 47held outn = 47Mismatch was a deleted last paragraph; writer bug fixed after the run, test set not re-run.
  • Model cost per contract (median, list price)US$0.0089held out16 calls, 22.9k prompt + 1.5k output tokens

Dataset

CUAD v1 (510 SEC EDGAR contracts, expert clause labels, CC BY 4.0): 40 train-split dev contracts for tuning, 30 test-split contracts run once after the configuration was frozen. Playbook and precedents are synthetic; planted deviations from a clause-template bank.

Caveats

  • The playbook and precedent library are synthetic, written for this demo; only clause presence and location come from CUAD.
  • Per-rule samples are small (3 to 6 labelled contracts per rule): warranty duration found in 0 of 3 and insurance in 2 of 5.
  • Only 17 of 30 test contracts had a buyer side, so only those could be planted.
  • Not measured: real firm playbooks, Microsoft Word itself, agreement with a lawyer's own redline, latency on a dedicated card.
  • One writer bug found by the test run was fixed afterwards; the test set was not re-run.

Nightly smoke check

Loading the nightly status…

Result
pass
Run
25 Sep 2026
Latency, this run
n/a
p50 over passed runs
28 s
Receipts
20
Model calls
n/a
Tokens
n/a
Cost per run
$0.007

Self-host verification

Verified on 25 Sep 2026: Fresh clone of the branch into a clean directory, docker build of docker/api/Dockerfile, the assemble prompt's api service with a named volume, pointed at the already-running local vLLM (Qwen3.8-27B NVFP4 on 127.0.0.1:8114) through host networking; then torn down.

Verified on 2026-09-25: the image builds, the service starts healthy, the Northwind sample passes end to end (9.7 s, 20 attested calls, 7 tracked changes, 0 untraced insertions, the file's checks pass), the signed record verifies and fails when one finding is changed, and the .docx export downloads. The model server's own startup was not re-verified (no new GPU load).

Rehearsal bundle: legal-drafting-editor.zip (11 KB, 12 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.

Known limits

  • The playbook and precedent library in the demo are synthetic; a firm's own need to be loaded (as JSON) and were not tested.
  • Paragraphs with fields, hyperlinks, drawings or earlier tracked changes get a comment, not an in-place edit; inserted clauses are not numbered.
  • Checked against the OOXML rules and LibreOffice; not yet opened in Microsoft Word by us.
  • Position errors: 2% of planted clauses got the wrong one of five positions (v2 held-out run) and no planted deviation was missed; like-for-like clause detection against CUAD is 0.77 / 0.79, and warranty-duration and insurance clauses are the ones most often missed.
  • The client-side playbook needs a buyer side: contracts between equals (joint ventures, cooperation agreements) get findings but party-specific precedents stay in comments.

Models and licences

Standard tier. Licence posture: permissive (Apache, MIT or BSD).

  • Editor: DOCX reading and the tracked-change writer, candidate search, traceability check, file checks, signed record and ledger (no model; CPU)decosa-api drafting module (decosa_api/verticals/drafting)AGPL-3.0-or-later
  • Model: places each clause against the playbook, proposes the find-and-replace edits, judges the grounding of each changed sentence, re-reads the edited clauseQwen3.8-27B (NVFP4)Apache-2.0

All quality evidence

Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.

Lite · one 32 GB card, self-hosted (1)
  • Findings and redlines on CUAD: not measured separately: the same weights and prompts as the standard tier; speed on a 5090 not measurednot measured yet
Standard · one 96 GB card (measured; hosted demo) (9)
  • Clause detection vs CUAD expert labels, 8 units × 30 held-out test contracts: precision / recall: 0.82 / 0.85 under the v2 mapping (v1: 0.81 / 0.79, re-scored); like for like with v1's mapping 0.77 / 0.79 (v1: 0.78 / 0.82), so no gain theredocs/evals/legal-drafting-editor.md (v2). The v2 gain is the new CONSEQ rule, which finds the exclusions of indirect damages CUAD files under "Cap On Liability" (9 of 9); warranty duration 0 of 3 and insurance 2 of 5 labelled contracts found
  • Planted deviations flagged (102 planted clauses, 17 CUAD test contracts): precision / recall: 0.98 / 1.00 (52 of 52 deviations flagged, 1 false flag); exact position (of 5) 0.98docs/evals/legal-drafting-editor.md; v1 was 0.89 / 0.95, exact 0.85, on 13 contracts, and 0.95 / 0.97, exact 0.92, once 6 items a harness bug had left out of the file are removed
  • Contracts where the buyer side was named (needed for party-specific precedents): 17 of 30 (v1: 13); both names among CUAD's labelled parties in 17 of 17docs/evals/legal-drafting-editor.md
  • Redlines that are valid, reject-all = input, accept-all = proposal: 47 of 47docs/evals/legal-drafting-editor.md, every output file of the v2 held-out run
  • Redlines LibreOffice opens and whose Accept All / Reject All match: 47 of 47 open, Reject All 47 of 47, Accept All 46 of 47; 284 of 284 comments importeddocs/evals/legal-drafting-editor.md: the mismatch was a deleted last paragraph (LibreOffice and Word keep an empty one); fixed after the run and tested, not re-run on the test set
  • Inserted text not traceable to a firm precedent: 0 in 136 tracked changesdocs/evals/legal-drafting-editor.md; enforced in code: an untraced chunk sends the edit back to the precedent's sentence
  • Targeted edits the checks sent back to the precedent's own sentence: 26 of 77docs/evals/legal-drafting-editor.md (v2 held-out run)
  • Model cost per contract, eleven rules (median): US$0.0089 at list price ($0.30 / $1.50 per million tokens), 16 calls (v1: US$0.0065, 14 calls)docs/evals/legal-drafting-editor.md, 47 runs
  • Levers tried on 40 dev contracts and left off: Self-consistency (3 readings): 2.6x the cost, no detection gain. Topic gate: +0.05 precision, -0.06 recall. Embedding retrieval (bge-small, CPU): +0 to +1 point of labelled text shown. Log-probabilities: the gateway does not return them.docs/evals/legal-drafting-editor.md, lever study
Wanted · a larger second reader on your own hardware (1)
  • CUAD precision and recall on the same 8 rules: not measured yet

How we measure · All tools