Eval: privileged drafting editor (use case 33)
v2 (25 Sep 2026, on a pre-release build)
Same model and route as v1: Qwen3.8-27B through the model gateway (every call receipted), temperature 0, thinking off, on
a GPU and gateway shared with other workloads' work. At the coordinator's request, all v2 runs used at most 2 requests in
flight (DECOSA_LLM_CONCURRENCY=2, one contract at a time), so latency differs from v1 for that reason alone.
Protocol
- Dev: 40 contracts drawn with a fixed seed from CUAD's train split (up to 60,000 characters; the two hosted samples excluded). Every lever was tried and chosen here. Exemplar pool: the other 366 train contracts, the only source of few-shot exemplars. Test: the same 30 held-out test-split contracts as v1, run once, after the configuration below was frozen and committed (249aa41). No test contract was read while tuning.
- One thing changed after the test run: a DOCX writer bug that run found (below). The fix changes no finding or position, and the test set was not re-run.
- Artifacts:
docs/evals/legal-drafting-editor/v2/(results.jsonl,summary.json,dev-levers.json,test-v1-vs-v2.json,v1-planting-audit.json). The v1 files are unchanged.
What changed
- Rule and label alignment. CUAD's "Cap On Liability" covers the cap amount, exclusions of indirect or
consequential damages, sole-remedy clauses and time limits on claims. The v1 rule "Limitation of liability" asked
only about the amount. v2 splits it into CAP (the amount, unchanged in substance) and a new CONSEQ rule
(exclusion of indirect and consequential damages; comment-only, precedent P-CONSEQ-01). The rule-to-CUAD mapping is
now written down (
MAPPING_V1and each rule'scuadfield): CUAD "Cap On Liability" is matched by CAP or CONSEQ, and NONCOMP, whose title has always said "non-compete and exclusivity", is matched by CUAD "Non-Compete" or "Exclusivity". Every detection number is given under both mappings. The v1 mapping is the like-for-like comparison; the v2 mapping is fairer to what the rules say, but part of its gain is the redefinition itself, not better reading. - Rule questions for IP, TERM, WARR and NONCOMP now say what is not their clause (keeping existing IP, renewal notices, warranties with no stated period, "non-exclusive" statements). The finding prompt says a clause on topic that binds only the Counterparty is not absent.
- Few-shot exemplars from the exemplar pool: per rule topic, up to two labelled clauses that count and two near misses that do not (the other half of "Cap On Liability"; confusable CUAD categories such as renewal notices for TERM, licence grants for IP, no-hire clauses for NONCOMP), picked by BM25 closeness to the candidate paragraphs. A first version took unlabelled search hits as near misses; hand-checking showed real clauses among them (CUAD's labels are incomplete), so only labelled spans are used.
- Retrieval: a per-rule
candidatescount; 8 instead of 6 for CAP, CONSEQ, IP, NONCOMP and INS. - Party detection: more defined-term styles ("Supplier" or "Fonterra"; hereinafter collectively "ENVISION"; bare capitals such as (MICOA); company short names such as Lucid Inc. -> Lucid), and the parties prompt counts alliances, sponsorships and partnerships with a recipient side as having a Client.
- Planting harness bug (eval code, v1 too). When two rules' CUAD labels pointed at the same paragraph (EDGAR text
often holds several clauses on one line), the second planted clause overwrote the first, or removed it, so the first
rule's truth was not in the contract the model read. v2 never overwrites or removes a planted paragraph (asserted).
Rebuilt from the v1 test results without new model calls (
scripts/drafting/audit_v1_planting.py): 6 of v1's 78 test items were affected, including one of its two "missed deviations" (a CAP clause that was not in the file). - Available but off by default: self-consistency (
DECOSA_DRAFTING_SAMPLES, 3 readings, majority, any disagreement goes to review) and a "holds_clause" topic gate. Both are explained in the lever study.
Held-out results (30 test contracts, run once): v1 vs v2
| Measure | v1 (as published) | v2 |
|---|---|---|
| Clause detection, v1 mapping (8 rules, like for like): precision / recall | 0.78 / 0.82 | 0.77 / 0.79 |
| Clause detection, v2 mapping (8 CUAD units; CAP or CONSEQ for "Cap On Liability", NONCOMP for "Non-Compete" or "Exclusivity") | 0.81 / 0.79 (v1's own findings, re-scored) | 0.82 / 0.85 |
| Located in the labelled paragraph, when both found | 60 / 62 | 67 / 70 |
| Contracts with a buyer side named, so planting could run | 13 / 30 | 17 / 30 (both names among CUAD's labelled party spans in 17 / 17) |
| Planted deviations flagged, six v1 rules: precision / recall | 0.89 / 0.95 (78 items); 0.95 / 0.97 without the 6 harness-broken items | 0.98 / 1.00 (102 items; 52 of 52 deviations flagged, 1 false flag) |
| Exact position of a planted clause (5 classes) | 0.85; 0.92 without the broken items | 0.98 (100 / 102) |
| Same 12 contracts both versions planted: flag P / R, exact | 0.90 / 0.95, 0.88 (as run) | 0.97 / 1.00, 0.97 |
| Redlines valid, reject-all = input, accept-all = proposal (our checks) | 43 / 43 | 47 / 47 |
| LibreOffice opens them; its Accept All / Reject All give our texts | 43 / 43 | 47 / 47 open; Reject All 47 / 47; Accept All 46 / 47 (bug below, fixed); comments 284 / 284 |
| Inserted text not traceable to a precedent | 0 in 117 tracked changes | 0 in 136 |
| Targeted edits sent back to the precedent's own sentence by a check | 19 of 59 | 26 of 77 (model wording never untraced; grounding on tracked changes: 47 supported, 14 partial, 3 unsupported, 3 contradicted) |
| Signed records that verify | 43 / 43 | 47 / 47 |
| Per contract (median of runs) | 14 calls, US$0.0065, 16.3k prompt + 1.3k output tokens | 16 calls, US$0.0089, 22.9k prompt + 1.5k output tokens (+36%) |
Prices are the gateway's list price ($0.30 / $1.50 per million tokens). Whole v2 test run: 746 calls, US$0.43.
Per unit, clause detection (precision / recall), test:
| Unit (CUAD category) | v1 mapping: v1 -> v2 | v2 mapping: v1 -> v2 | Note |
|---|---|---|---|
| Governing law | 1.00 / 1.00 -> 1.00 / 1.00 | same | |
| Anti-assignment | 1.00 / 0.94 -> 1.00 / 0.94 | same | |
| Cap On Liability | CAP 1.00 / 0.33 -> 1.00 / 0.22 | 1.00 / 0.33 -> 1.00 / 1.00 | CUAD labels this category in 9 of the 30 contracts. CAP alone finds a cap in 2 of them (v1: 3); CONSEQ alone finds an exclusion in all 9, with no false find. This is the alignment, not better reading of caps. |
| Non-compete (+ exclusivity) | 0.42 / 0.71 -> 0.47 / 1.00 | 0.67 / 0.62 -> 0.67 / 0.77 | |
| Termination for convenience | 0.40 / 0.80 -> 0.50 / 0.80 | same | 4 false finds left (5 labelled contracts) |
| IP ownership | 0.67 / 1.00 -> 0.56 / 0.83 | same | 6 labelled contracts; CUAD counts only assignment of new IP |
| Insurance | 1.00 / 0.60 -> 1.00 / 0.40 | same | 5 labelled contracts: one miss more |
| Warranty duration | 0.33 / 0.33 -> 0.00 / 0.00 | same | 3 labelled contracts, all missed; 2 false finds |
Planted, per rule (flag precision / recall, exact): governing law 1.00 / 1.00, 17/17; CAP 1.00 / 1.00, 17/17; warranty 1.00 / 1.00, 17/17; assignment 1.00 / 1.00, 17/17; insurance 1.00 / 1.00, 17/17; non-compete 0.83 / 1.00, 15/17 (one compliant clause read as walk-away, one preferred read as fallback). v1: governing law 0.73 / 1.00, CAP 1.00 / 0.86, insurance 1.00 / 0.80, non-compete 0.67 / 1.00.
Reading the v2 numbers
- Where v2 is better: positions. On planted clauses, exact position went from 0.85 (0.92 once v1's harness bug is taken out) to 0.98, and no deviation was missed (52 of 52). Four more contracts could be planted (17 vs 13), and the named parties were right by CUAD's party labels in all 17.
- Where it is not: like-for-like clause detection did not improve (0.78 / 0.82 -> 0.77 / 0.79). The v2-mapping gain (0.79 -> 0.85 recall) comes from the new CONSEQ rule matching the exclusions that CUAD files under "Cap On Liability": that is a better product (the exclusion now gets a finding) and a fairer mapping, but not better reading of the rules v1 had. Insurance lost one of 5 labelled contracts and warranty duration all 3 (it had 1 of 3); at these sizes (3 to 6 labelled contracts per rule) one contract moves recall by 0.2 to 0.3, so these are noise-sized, but they are not improvements.
- The biggest part of v1's "planted" error was the harness. 6 of 78 v1 items had no planted clause in the file. Without them v1 was 0.95 / 0.97, exact 0.92; v2 is 0.98 / 1.00, exact 0.98. The fix matters for reading both.
- Cost: +36% per contract (US$0.0065 -> US$0.0089): one more rule (CONSEQ), the exemplars (about +2,300 prompt tokens per run) and more candidates for five rules.
Lever study (dev, 40 contracts; v2/dev-levers.json)
Each row adds one change to the row before it unless noted. "v1ref" is v1's code with the fixed planting harness (the same dev contracts). Detection is P / R under the v2 mapping (about 135 labelled pairs, so ±0.03 is 4 pairs and within noise); planted is flag P / R and exact position on the six v1 rules. Cost is the median per run at list price.
| Run | Change | Detection v2 (v1 mapping) | Planted P / R, exact | Planted contracts | Median cost | Prompt tokens / run |
|---|---|---|---|---|---|---|
| v1ref | v1 code, fixed harness | 0.73 / 0.82 (0.71 / 0.84) | 0.99 / 0.97, 0.93 | 22 | $0.0072 | 17.7k |
| A | + CAP / CONSEQ split, NONCOMP scope, prompt clarifications, party detection | 0.74 / 0.92 (0.70 / 0.87) | 0.98 / 0.99, 0.97 | 32 | $0.0086 | 20.6k |
| B | + rule questions say what is not the clause | 0.78 / 0.90 (0.74 / 0.85) | 0.99 / 0.98, 0.98 | 31 | $0.0084 | 20.9k |
| C | B + exemplars | 0.76 / 0.91 | 1.00 / 0.98, 0.98 | 32 | $0.0092 | 23.2k |
| G | B + 8 candidates for five rules | 0.77 / 0.91 | 1.00 / 0.97, 0.98 | 30 | $0.0088 | 22.0k |
| D (shipped) | B + exemplars + 8 candidates | 0.77 / 0.94 (0.71 / 0.86) | 0.99 / 0.97, 0.98 | 31 | $0.0097 | 24.6k |
| E | D + topic gate ("holds_clause" first) | 0.82 / 0.88 (0.78 / 0.83) | 0.98 / 0.98, 0.98 | 31 | $0.0096 | 25.4k |
| E2 | E without exemplars | 0.81 / 0.89 | 0.99 / 0.95, 0.97 | 31 | $0.0090 | 23.0k |
| F | E + self-consistency, 3 readings | 0.83 / 0.88 | 1.00 / 0.98, 0.98; with review counted as a flag 0.97 / 0.99 | 31 | $0.0247 | 68.1k |
- Alignment and party detection (A) carried most of the gain: "Cap On Liability" recall 0.41 -> 0.86 through CONSEQ, and 10 more planted contracts. Rule wording (B) cut false finds for termination for convenience (0.70 -> 0.84 precision). Both are almost free.
- Exemplars and more candidates (C, G, D): each alone is within noise of B; together they gave the best recall (0.94; NONCOMP 0.94 recall, "Cap On Liability" 0.91) for about +15% cost. They are shipped because a missed clause is the costlier error, but the evidence is thin: on test, detection did not rise under the v1 mapping.
- Topic gate (E) traded recall for precision (0.77 / 0.94 -> 0.82 / 0.88), mostly by dropping "Cap On Liability" finds; same F1. Off by default for the same reason.
- Self-consistency (F): 2.6x the cost for no detection change and one more planted deviation reached review.
Available (
DECOSA_DRAFTING_SAMPLES=3), off by default. - Log-probabilities for routing low-confidence calls to review: not possible on the hosted route. Checked 25 Sep:
the gateway drops
logprobs(a direct request withlogprobs: truereturned none), andDECOSA_JUDGMENT_GATEWAY_LOGPROBSis unset. Self-consistency is the stand-in. - Hybrid lexical + embedding retrieval (not shipped). Measured without model calls
(
scripts/drafting/retrieval_diag.py): the share of CUAD-labelled spans the model is shown. Lexical search with 6 candidates shows a labelled span for 134 of 138 contract-rule pairs and 82% of all labelled spans; 8 candidates: 85%; 10: 86%; heading-aware indexing: 82%. Adding bge-small-en-v1.5 (CPU, 25 ms per window) scores at weight 0.6: 83% at 6 candidates, 86% at 8. That is +0 to +1 point for a new model dependency, so only the candidate count shipped. A bge-m3 run was stopped because it competed for CPU with other jobs on the box.
Bug found by the held-out run (fixed after it)
- Deleting the document's last paragraph. In one planted contract the non-compete was the body's last paragraph
and the walk-away action deleted it. Our writer marked its paragraph mark as deleted; Word and LibreOffice cannot
remove the last paragraph mark, so LibreOffice's Accept All kept an empty paragraph where our own check expected none
(46 of 47 matched). The writer now deletes only the text of a paragraph with no paragraph after it; the accept-all
view and the file check expect the empty paragraph. Test:
test_deleting_the_last_paragraph_keeps_an_empty_one; checked with LibreOffice's own Accept All / Reject All on a synthetic file. v1 had the same bug; it did not come up in its 43 files.
Not measured
Real firm playbooks and precedent libraries; Microsoft Word itself; agreement with a lawyer's own redline; whether the exemplars help a firm's own custom rules (they only apply to the default rule topics); latency on a dedicated card (v2 runs used 2 requests in flight on a saturated shared GPU).
Reproduce (v2)
export DECOSA_LLM_URL=<your model endpoint> DECOSA_LLM_KEY=<your key>; export DECOSA_LLM_CONCURRENCY=2
.venv/bin/python scripts/drafting/build_exemplars.py # data/exemplars.json from the train split
.venv/bin/python scripts/drafting_eval.py --split dev --n 40 --parallel 1 --out DIR [--no-exemplars] [--candidates 6] [--gate] [--samples 3]
.venv/bin/python scripts/drafting/compare_runs.py NAME=DIR ... # lever table
.venv/bin/python scripts/drafting/retrieval_diag.py [--k 8] [--headings] # retrieval recall, no model calls
.venv/bin/python scripts/drafting_eval.py --split test --n 30 --parallel 1 --lo --out DIR # held-out, once
v1 (first release, 25 Sep 2026)
Run 25 Sep 2026 on our server. Model: Qwen3.8-27B through the model gateway (receipted), temperature 0, thinking off, on a
GPU and gateway shared with other workloads' work. Script: scripts/drafting_eval.py; raw results:
docs/evals/legal-drafting-editor/results.jsonl, summary: summary.json (test) and dev-summary.json (dev).
Data
- Contracts: CUAD v1 (The Atticus Project, CC BY 4.0): 510 commercial contracts from SEC EDGAR filings with clause labels by trained annotators. Test: 30 contracts drawn with seed 33 from CUAD's own test split (102 contracts), limited to 60,000 characters each for cost. Dev: 6 contracts from the train split (excluding the two hosted samples), used to develop the prompts and the search. The test set was run once, after the prompts were frozen.
- Playbook and precedents: synthetic, written for this demo (a fictional firm, 10 rules, 13 precedent clauses; see
decosa_api/verticals/drafting/data/). CUAD has no playbook, so the playbook positions are ours; only clause presence and location come from CUAD's labels. - Each contract as a DOCX made from CUAD's text (one paragraph per line), so the runs test our own DOCX writer on plain paragraphs. Real firm DOCX files (styles, numbering, fields) were only tested in unit tests (bold runs, bookmarks, hyperlinks, existing comments and revisions), not at scale.
What was measured
Two full runs per contract, all ten rules:
- As filed: clause detection. For the 8 rules that map to a CUAD category, "present" is our finding placing the clause anywhere but "absent"; truth is a non-empty CUAD label. "Located" = one of our paragraphs overlaps a labelled span.
- Planted: six rules planted per contract from a bank of clause templates (governing law, liability cap, warranty, assignment, insurance, non-compete) at a position drawn with a fixed seed: preferred, fallback, outside, walk-away, and for governing law also removed. CUAD's clauses for those categories are replaced, so the planted clause is the truth. "Flagged" = placed outside or at walk-away, or reported missing where the playbook acts on a missing clause. Planting needs the two parties' defined terms: 13 of 30 contracts had them (the other 17 are joint filings, cooperation and non-compete agreements with no buyer side, or text with no quoted defined terms).
- Every output file (43): package validation, reject-all equals the input, accept-all equals the proposal, and
LibreOffice 7.4 opening it headless and applying its own Accept All and Reject All (UNO) in the
decosa-lo-validateimage (scripts/drafting/Dockerfile.lo,lo_roundtrip.py). LibreOffice aborts when a DOCX with comments is loaded through UNO headless (7.4 and 25.2 both), so that step runs on a copy with the comment anchors removed; comments are counted from a plainsoffice --convert-to odtof the real file. - Traceability: every inserted chunk in every tracked change checked against the precedent it cites.
Results (test, 30 contracts)
| Measure | Result |
|---|---|
| Clause detection vs CUAD labels, 8 rules × 30 contracts | precision 0.78, recall 0.82 (62 TP, 18 FP, 14 FN, 146 TN); located in 60 of 62 |
| Planted deviations flagged (78 planted clauses, 13 contracts) | precision 0.89, recall 0.95 (39 TP, 5 FP, 2 FN, 32 TN) |
| Exact position on planted clauses (5 classes) | 0.85; the planted paragraph found in 66 of 68 |
| Redlines valid, reject-all = input, accept-all = proposal | 43 / 43 each |
| LibreOffice opens them; its Accept All and Reject All give our texts | 43 / 43 each; 250 of 250 comments imported |
| Inserted text not traceable to a precedent, in the output | 0 in 117 tracked changes |
| Targeted edits sent back to the precedent's own sentence by a check | 19 of 59 (7 not correct English on re-read, 6 grounding contradicted or unsupported, 5 the model asked for it, 1 find string not unique) |
| Signed records that verify | 43 / 43 |
| Per contract (median) | 14 model calls, 12 s, US$0.0065 at list price ($0.30 / $1.50 per M tokens) |
Per rule, detection (precision / recall): governing law 1.00 / 1.00; assignment 1.00 / 0.94; insurance 1.00 / 0.60; liability cap 1.00 / 0.33; IP ownership 0.67 / 1.00; termination for convenience 0.40 / 0.80; non-compete 0.42 / 0.71; warranty duration 0.33 / 0.33 (only 3 labelled contracts). Planted, per rule (flag precision / recall): governing law 0.73 / 1.00, liability cap 1.00 / 0.86, warranty 1.00 / 1.00, assignment 1.00 / 1.00, insurance 1.00 / 0.80, non-compete 0.67 / 1.00.
Reading the numbers
- Definitions differ from CUAD's. CUAD's "Cap on Liability" includes exclusions of consequential damages and time limits on claims; our rule is about the amount of the cap, so the model calls a bare exclusion "absent" (recall 0.33). CUAD's "Non-Compete" covers either party; our rule is about restrictions on the client, and the model also reports exclusivity clauses, which CUAD labels separately. These are disagreements of definition as much as errors.
- The failure that matters most is a missed deviation (the lawyer is not warned): 2 of 41 planted deviations were missed (one liability cap, one insurance limit). False flags (5: three compliant governing-law clauses, two non-competes) cost a lawyer a look, not a risk.
- "Never invented" is enforced in code, not trusted to the model. The model's own wording was never untraced in this run, but the check is what guarantees it: an untraced chunk sends the edit back to the precedent's sentence.
- Minimal edits: a targeted edit changes a median 41% of its paragraph's words (planted clauses are one short sentence, so this is high); no paragraph outside the edits changed in any file (accept-all check).
- Not measured: real firm playbooks and precedent libraries; DOCX files with heavy numbering, fields and tables at scale; Microsoft Word itself (we checked the OOXML rules Word enforces and used LibreOffice as the independent reader); agreement with a lawyer's own redline; latency on a dedicated card.
Changes after the test run
One: the party-name pattern now also reads single-quoted defined terms (('the Customer')), found in contracts the
planting step skipped. It changes which contracts can be planted, not any number above, and was not re-scored on test.
Reproduce
export DECOSA_LLM_URL=<your model endpoint> DECOSA_LLM_KEY=<your key>;
docker build -f scripts/drafting/Dockerfile.lo -t decosa-lo-validate:bookworm scripts/drafting
.venv/bin/python scripts/drafting_eval.py --split test --n 30 --parallel 3 --lo # CUAD in ~/data/cuad (data.zip from the CUAD GitHub repo)