Skip to content
decosa

The full write-up behind the numbers on the tool’s page, as the team that built it wrote it: data, method, results and caveats. Internal names are removed; nothing else is edited.

Eval: privileged drafting editor (use case 33)

v2 (25 Sep 2026, on a pre-release build)

Same model and route as v1: Qwen3.8-27B through the model gateway (every call receipted), temperature 0, thinking off, on a GPU and gateway shared with other workloads' work. At the coordinator's request, all v2 runs used at most 2 requests in flight (DECOSA_LLM_CONCURRENCY=2, one contract at a time), so latency differs from v1 for that reason alone.

Protocol

  • Dev: 40 contracts drawn with a fixed seed from CUAD's train split (up to 60,000 characters; the two hosted samples excluded). Every lever was tried and chosen here. Exemplar pool: the other 366 train contracts, the only source of few-shot exemplars. Test: the same 30 held-out test-split contracts as v1, run once, after the configuration below was frozen and committed (249aa41). No test contract was read while tuning.
  • One thing changed after the test run: a DOCX writer bug that run found (below). The fix changes no finding or position, and the test set was not re-run.
  • Artifacts: docs/evals/legal-drafting-editor/v2/ (results.jsonl, summary.json, dev-levers.json, test-v1-vs-v2.json, v1-planting-audit.json). The v1 files are unchanged.

What changed

  1. Rule and label alignment. CUAD's "Cap On Liability" covers the cap amount, exclusions of indirect or consequential damages, sole-remedy clauses and time limits on claims. The v1 rule "Limitation of liability" asked only about the amount. v2 splits it into CAP (the amount, unchanged in substance) and a new CONSEQ rule (exclusion of indirect and consequential damages; comment-only, precedent P-CONSEQ-01). The rule-to-CUAD mapping is now written down (MAPPING_V1 and each rule's cuad field): CUAD "Cap On Liability" is matched by CAP or CONSEQ, and NONCOMP, whose title has always said "non-compete and exclusivity", is matched by CUAD "Non-Compete" or "Exclusivity". Every detection number is given under both mappings. The v1 mapping is the like-for-like comparison; the v2 mapping is fairer to what the rules say, but part of its gain is the redefinition itself, not better reading.
  2. Rule questions for IP, TERM, WARR and NONCOMP now say what is not their clause (keeping existing IP, renewal notices, warranties with no stated period, "non-exclusive" statements). The finding prompt says a clause on topic that binds only the Counterparty is not absent.
  3. Few-shot exemplars from the exemplar pool: per rule topic, up to two labelled clauses that count and two near misses that do not (the other half of "Cap On Liability"; confusable CUAD categories such as renewal notices for TERM, licence grants for IP, no-hire clauses for NONCOMP), picked by BM25 closeness to the candidate paragraphs. A first version took unlabelled search hits as near misses; hand-checking showed real clauses among them (CUAD's labels are incomplete), so only labelled spans are used.
  4. Retrieval: a per-rule candidates count; 8 instead of 6 for CAP, CONSEQ, IP, NONCOMP and INS.
  5. Party detection: more defined-term styles ("Supplier" or "Fonterra"; hereinafter collectively "ENVISION"; bare capitals such as (MICOA); company short names such as Lucid Inc. -> Lucid), and the parties prompt counts alliances, sponsorships and partnerships with a recipient side as having a Client.
  6. Planting harness bug (eval code, v1 too). When two rules' CUAD labels pointed at the same paragraph (EDGAR text often holds several clauses on one line), the second planted clause overwrote the first, or removed it, so the first rule's truth was not in the contract the model read. v2 never overwrites or removes a planted paragraph (asserted). Rebuilt from the v1 test results without new model calls (scripts/drafting/audit_v1_planting.py): 6 of v1's 78 test items were affected, including one of its two "missed deviations" (a CAP clause that was not in the file).
  7. Available but off by default: self-consistency (DECOSA_DRAFTING_SAMPLES, 3 readings, majority, any disagreement goes to review) and a "holds_clause" topic gate. Both are explained in the lever study.

Held-out results (30 test contracts, run once): v1 vs v2

Measure v1 (as published) v2
Clause detection, v1 mapping (8 rules, like for like): precision / recall 0.78 / 0.82 0.77 / 0.79
Clause detection, v2 mapping (8 CUAD units; CAP or CONSEQ for "Cap On Liability", NONCOMP for "Non-Compete" or "Exclusivity") 0.81 / 0.79 (v1's own findings, re-scored) 0.82 / 0.85
Located in the labelled paragraph, when both found 60 / 62 67 / 70
Contracts with a buyer side named, so planting could run 13 / 30 17 / 30 (both names among CUAD's labelled party spans in 17 / 17)
Planted deviations flagged, six v1 rules: precision / recall 0.89 / 0.95 (78 items); 0.95 / 0.97 without the 6 harness-broken items 0.98 / 1.00 (102 items; 52 of 52 deviations flagged, 1 false flag)
Exact position of a planted clause (5 classes) 0.85; 0.92 without the broken items 0.98 (100 / 102)
Same 12 contracts both versions planted: flag P / R, exact 0.90 / 0.95, 0.88 (as run) 0.97 / 1.00, 0.97
Redlines valid, reject-all = input, accept-all = proposal (our checks) 43 / 43 47 / 47
LibreOffice opens them; its Accept All / Reject All give our texts 43 / 43 47 / 47 open; Reject All 47 / 47; Accept All 46 / 47 (bug below, fixed); comments 284 / 284
Inserted text not traceable to a precedent 0 in 117 tracked changes 0 in 136
Targeted edits sent back to the precedent's own sentence by a check 19 of 59 26 of 77 (model wording never untraced; grounding on tracked changes: 47 supported, 14 partial, 3 unsupported, 3 contradicted)
Signed records that verify 43 / 43 47 / 47
Per contract (median of runs) 14 calls, US$0.0065, 16.3k prompt + 1.3k output tokens 16 calls, US$0.0089, 22.9k prompt + 1.5k output tokens (+36%)

Prices are the gateway's list price ($0.30 / $1.50 per million tokens). Whole v2 test run: 746 calls, US$0.43.

Per unit, clause detection (precision / recall), test:

Unit (CUAD category) v1 mapping: v1 -> v2 v2 mapping: v1 -> v2 Note
Governing law 1.00 / 1.00 -> 1.00 / 1.00 same
Anti-assignment 1.00 / 0.94 -> 1.00 / 0.94 same
Cap On Liability CAP 1.00 / 0.33 -> 1.00 / 0.22 1.00 / 0.33 -> 1.00 / 1.00 CUAD labels this category in 9 of the 30 contracts. CAP alone finds a cap in 2 of them (v1: 3); CONSEQ alone finds an exclusion in all 9, with no false find. This is the alignment, not better reading of caps.
Non-compete (+ exclusivity) 0.42 / 0.71 -> 0.47 / 1.00 0.67 / 0.62 -> 0.67 / 0.77
Termination for convenience 0.40 / 0.80 -> 0.50 / 0.80 same 4 false finds left (5 labelled contracts)
IP ownership 0.67 / 1.00 -> 0.56 / 0.83 same 6 labelled contracts; CUAD counts only assignment of new IP
Insurance 1.00 / 0.60 -> 1.00 / 0.40 same 5 labelled contracts: one miss more
Warranty duration 0.33 / 0.33 -> 0.00 / 0.00 same 3 labelled contracts, all missed; 2 false finds

Planted, per rule (flag precision / recall, exact): governing law 1.00 / 1.00, 17/17; CAP 1.00 / 1.00, 17/17; warranty 1.00 / 1.00, 17/17; assignment 1.00 / 1.00, 17/17; insurance 1.00 / 1.00, 17/17; non-compete 0.83 / 1.00, 15/17 (one compliant clause read as walk-away, one preferred read as fallback). v1: governing law 0.73 / 1.00, CAP 1.00 / 0.86, insurance 1.00 / 0.80, non-compete 0.67 / 1.00.

Reading the v2 numbers

  • Where v2 is better: positions. On planted clauses, exact position went from 0.85 (0.92 once v1's harness bug is taken out) to 0.98, and no deviation was missed (52 of 52). Four more contracts could be planted (17 vs 13), and the named parties were right by CUAD's party labels in all 17.
  • Where it is not: like-for-like clause detection did not improve (0.78 / 0.82 -> 0.77 / 0.79). The v2-mapping gain (0.79 -> 0.85 recall) comes from the new CONSEQ rule matching the exclusions that CUAD files under "Cap On Liability": that is a better product (the exclusion now gets a finding) and a fairer mapping, but not better reading of the rules v1 had. Insurance lost one of 5 labelled contracts and warranty duration all 3 (it had 1 of 3); at these sizes (3 to 6 labelled contracts per rule) one contract moves recall by 0.2 to 0.3, so these are noise-sized, but they are not improvements.
  • The biggest part of v1's "planted" error was the harness. 6 of 78 v1 items had no planted clause in the file. Without them v1 was 0.95 / 0.97, exact 0.92; v2 is 0.98 / 1.00, exact 0.98. The fix matters for reading both.
  • Cost: +36% per contract (US$0.0065 -> US$0.0089): one more rule (CONSEQ), the exemplars (about +2,300 prompt tokens per run) and more candidates for five rules.

Lever study (dev, 40 contracts; v2/dev-levers.json)

Each row adds one change to the row before it unless noted. "v1ref" is v1's code with the fixed planting harness (the same dev contracts). Detection is P / R under the v2 mapping (about 135 labelled pairs, so ±0.03 is 4 pairs and within noise); planted is flag P / R and exact position on the six v1 rules. Cost is the median per run at list price.

Run Change Detection v2 (v1 mapping) Planted P / R, exact Planted contracts Median cost Prompt tokens / run
v1ref v1 code, fixed harness 0.73 / 0.82 (0.71 / 0.84) 0.99 / 0.97, 0.93 22 $0.0072 17.7k
A + CAP / CONSEQ split, NONCOMP scope, prompt clarifications, party detection 0.74 / 0.92 (0.70 / 0.87) 0.98 / 0.99, 0.97 32 $0.0086 20.6k
B + rule questions say what is not the clause 0.78 / 0.90 (0.74 / 0.85) 0.99 / 0.98, 0.98 31 $0.0084 20.9k
C B + exemplars 0.76 / 0.91 1.00 / 0.98, 0.98 32 $0.0092 23.2k
G B + 8 candidates for five rules 0.77 / 0.91 1.00 / 0.97, 0.98 30 $0.0088 22.0k
D (shipped) B + exemplars + 8 candidates 0.77 / 0.94 (0.71 / 0.86) 0.99 / 0.97, 0.98 31 $0.0097 24.6k
E D + topic gate ("holds_clause" first) 0.82 / 0.88 (0.78 / 0.83) 0.98 / 0.98, 0.98 31 $0.0096 25.4k
E2 E without exemplars 0.81 / 0.89 0.99 / 0.95, 0.97 31 $0.0090 23.0k
F E + self-consistency, 3 readings 0.83 / 0.88 1.00 / 0.98, 0.98; with review counted as a flag 0.97 / 0.99 31 $0.0247 68.1k
  • Alignment and party detection (A) carried most of the gain: "Cap On Liability" recall 0.41 -> 0.86 through CONSEQ, and 10 more planted contracts. Rule wording (B) cut false finds for termination for convenience (0.70 -> 0.84 precision). Both are almost free.
  • Exemplars and more candidates (C, G, D): each alone is within noise of B; together they gave the best recall (0.94; NONCOMP 0.94 recall, "Cap On Liability" 0.91) for about +15% cost. They are shipped because a missed clause is the costlier error, but the evidence is thin: on test, detection did not rise under the v1 mapping.
  • Topic gate (E) traded recall for precision (0.77 / 0.94 -> 0.82 / 0.88), mostly by dropping "Cap On Liability" finds; same F1. Off by default for the same reason.
  • Self-consistency (F): 2.6x the cost for no detection change and one more planted deviation reached review. Available (DECOSA_DRAFTING_SAMPLES=3), off by default.
  • Log-probabilities for routing low-confidence calls to review: not possible on the hosted route. Checked 25 Sep: the gateway drops logprobs (a direct request with logprobs: true returned none), and DECOSA_JUDGMENT_GATEWAY_LOGPROBS is unset. Self-consistency is the stand-in.
  • Hybrid lexical + embedding retrieval (not shipped). Measured without model calls (scripts/drafting/retrieval_diag.py): the share of CUAD-labelled spans the model is shown. Lexical search with 6 candidates shows a labelled span for 134 of 138 contract-rule pairs and 82% of all labelled spans; 8 candidates: 85%; 10: 86%; heading-aware indexing: 82%. Adding bge-small-en-v1.5 (CPU, 25 ms per window) scores at weight 0.6: 83% at 6 candidates, 86% at 8. That is +0 to +1 point for a new model dependency, so only the candidate count shipped. A bge-m3 run was stopped because it competed for CPU with other jobs on the box.

Bug found by the held-out run (fixed after it)

  • Deleting the document's last paragraph. In one planted contract the non-compete was the body's last paragraph and the walk-away action deleted it. Our writer marked its paragraph mark as deleted; Word and LibreOffice cannot remove the last paragraph mark, so LibreOffice's Accept All kept an empty paragraph where our own check expected none (46 of 47 matched). The writer now deletes only the text of a paragraph with no paragraph after it; the accept-all view and the file check expect the empty paragraph. Test: test_deleting_the_last_paragraph_keeps_an_empty_one; checked with LibreOffice's own Accept All / Reject All on a synthetic file. v1 had the same bug; it did not come up in its 43 files.

Not measured

Real firm playbooks and precedent libraries; Microsoft Word itself; agreement with a lawyer's own redline; whether the exemplars help a firm's own custom rules (they only apply to the default rule topics); latency on a dedicated card (v2 runs used 2 requests in flight on a saturated shared GPU).

Reproduce (v2)

export DECOSA_LLM_URL=<your model endpoint> DECOSA_LLM_KEY=<your key>; export DECOSA_LLM_CONCURRENCY=2
.venv/bin/python scripts/drafting/build_exemplars.py                         # data/exemplars.json from the train split
.venv/bin/python scripts/drafting_eval.py --split dev --n 40 --parallel 1 --out DIR [--no-exemplars] [--candidates 6] [--gate] [--samples 3]
.venv/bin/python scripts/drafting/compare_runs.py NAME=DIR ...               # lever table
.venv/bin/python scripts/drafting/retrieval_diag.py [--k 8] [--headings]     # retrieval recall, no model calls
.venv/bin/python scripts/drafting_eval.py --split test --n 30 --parallel 1 --lo --out DIR   # held-out, once

v1 (first release, 25 Sep 2026)

Run 25 Sep 2026 on our server. Model: Qwen3.8-27B through the model gateway (receipted), temperature 0, thinking off, on a GPU and gateway shared with other workloads' work. Script: scripts/drafting_eval.py; raw results: docs/evals/legal-drafting-editor/results.jsonl, summary: summary.json (test) and dev-summary.json (dev).

Data

  • Contracts: CUAD v1 (The Atticus Project, CC BY 4.0): 510 commercial contracts from SEC EDGAR filings with clause labels by trained annotators. Test: 30 contracts drawn with seed 33 from CUAD's own test split (102 contracts), limited to 60,000 characters each for cost. Dev: 6 contracts from the train split (excluding the two hosted samples), used to develop the prompts and the search. The test set was run once, after the prompts were frozen.
  • Playbook and precedents: synthetic, written for this demo (a fictional firm, 10 rules, 13 precedent clauses; see decosa_api/verticals/drafting/data/). CUAD has no playbook, so the playbook positions are ours; only clause presence and location come from CUAD's labels.
  • Each contract as a DOCX made from CUAD's text (one paragraph per line), so the runs test our own DOCX writer on plain paragraphs. Real firm DOCX files (styles, numbering, fields) were only tested in unit tests (bold runs, bookmarks, hyperlinks, existing comments and revisions), not at scale.

What was measured

Two full runs per contract, all ten rules:

  1. As filed: clause detection. For the 8 rules that map to a CUAD category, "present" is our finding placing the clause anywhere but "absent"; truth is a non-empty CUAD label. "Located" = one of our paragraphs overlaps a labelled span.
  2. Planted: six rules planted per contract from a bank of clause templates (governing law, liability cap, warranty, assignment, insurance, non-compete) at a position drawn with a fixed seed: preferred, fallback, outside, walk-away, and for governing law also removed. CUAD's clauses for those categories are replaced, so the planted clause is the truth. "Flagged" = placed outside or at walk-away, or reported missing where the playbook acts on a missing clause. Planting needs the two parties' defined terms: 13 of 30 contracts had them (the other 17 are joint filings, cooperation and non-compete agreements with no buyer side, or text with no quoted defined terms).
  3. Every output file (43): package validation, reject-all equals the input, accept-all equals the proposal, and LibreOffice 7.4 opening it headless and applying its own Accept All and Reject All (UNO) in the decosa-lo-validate image (scripts/drafting/Dockerfile.lo, lo_roundtrip.py). LibreOffice aborts when a DOCX with comments is loaded through UNO headless (7.4 and 25.2 both), so that step runs on a copy with the comment anchors removed; comments are counted from a plain soffice --convert-to odt of the real file.
  4. Traceability: every inserted chunk in every tracked change checked against the precedent it cites.

Results (test, 30 contracts)

Measure Result
Clause detection vs CUAD labels, 8 rules × 30 contracts precision 0.78, recall 0.82 (62 TP, 18 FP, 14 FN, 146 TN); located in 60 of 62
Planted deviations flagged (78 planted clauses, 13 contracts) precision 0.89, recall 0.95 (39 TP, 5 FP, 2 FN, 32 TN)
Exact position on planted clauses (5 classes) 0.85; the planted paragraph found in 66 of 68
Redlines valid, reject-all = input, accept-all = proposal 43 / 43 each
LibreOffice opens them; its Accept All and Reject All give our texts 43 / 43 each; 250 of 250 comments imported
Inserted text not traceable to a precedent, in the output 0 in 117 tracked changes
Targeted edits sent back to the precedent's own sentence by a check 19 of 59 (7 not correct English on re-read, 6 grounding contradicted or unsupported, 5 the model asked for it, 1 find string not unique)
Signed records that verify 43 / 43
Per contract (median) 14 model calls, 12 s, US$0.0065 at list price ($0.30 / $1.50 per M tokens)

Per rule, detection (precision / recall): governing law 1.00 / 1.00; assignment 1.00 / 0.94; insurance 1.00 / 0.60; liability cap 1.00 / 0.33; IP ownership 0.67 / 1.00; termination for convenience 0.40 / 0.80; non-compete 0.42 / 0.71; warranty duration 0.33 / 0.33 (only 3 labelled contracts). Planted, per rule (flag precision / recall): governing law 0.73 / 1.00, liability cap 1.00 / 0.86, warranty 1.00 / 1.00, assignment 1.00 / 1.00, insurance 1.00 / 0.80, non-compete 0.67 / 1.00.

Reading the numbers

  • Definitions differ from CUAD's. CUAD's "Cap on Liability" includes exclusions of consequential damages and time limits on claims; our rule is about the amount of the cap, so the model calls a bare exclusion "absent" (recall 0.33). CUAD's "Non-Compete" covers either party; our rule is about restrictions on the client, and the model also reports exclusivity clauses, which CUAD labels separately. These are disagreements of definition as much as errors.
  • The failure that matters most is a missed deviation (the lawyer is not warned): 2 of 41 planted deviations were missed (one liability cap, one insurance limit). False flags (5: three compliant governing-law clauses, two non-competes) cost a lawyer a look, not a risk.
  • "Never invented" is enforced in code, not trusted to the model. The model's own wording was never untraced in this run, but the check is what guarantees it: an untraced chunk sends the edit back to the precedent's sentence.
  • Minimal edits: a targeted edit changes a median 41% of its paragraph's words (planted clauses are one short sentence, so this is high); no paragraph outside the edits changed in any file (accept-all check).
  • Not measured: real firm playbooks and precedent libraries; DOCX files with heavy numbering, fields and tables at scale; Microsoft Word itself (we checked the OOXML rules Word enforces and used LibreOffice as the independent reader); agreement with a lawyer's own redline; latency on a dedicated card.

Changes after the test run

One: the party-name pattern now also reads single-quoted defined terms (('the Customer')), found in contracts the planting step skipped. It changes which contracts can be planted, not any number above, and was not re-scored on test.

Reproduce

export DECOSA_LLM_URL=<your model endpoint> DECOSA_LLM_KEY=<your key>;
docker build -f scripts/drafting/Dockerfile.lo -t decosa-lo-validate:bookworm scripts/drafting
.venv/bin/python scripts/drafting_eval.py --split test --n 30 --parallel 3 --lo    # CUAD in ~/data/cuad (data.zip from the CUAD GitHub repo)