Skip to content
decosa

The full write-up behind the numbers on the tool’s page, as the team that built it wrote it: data, method, results and caveats. Internal names are removed; nothing else is edited.

Privilege review and privilege log (30): eval

Run on 25 Sep 2026 on our server (one RTX PRO 6000, shared with other workloads' work the whole time) with scripts/privilege_eval.py. Model: Qwen3.8-27B NVFP4 (nvidia/Qwen3.8-27B-NVFP4 at revision 482ca0f3…, weights root 2fafb365…1091894), vLLM 0.29.0, thinking off. Prompts: review.py (PROMPTS_SHA256 a1bc0fef…). Metrics: docs/evals/privilege-log/metrics.json. Raw outputs are in ~/.cache/privilege-eval/ on our server (not committed).

Data and labels

Set What Size Labels
Enron Emails sampled from the mailboxes of Enron lawyers (Cash, Derrick, Haedicke, Sanders, Sager, Hodge, Taylor, Mann, Shackleton, Nemec, Kean, Steffes) in the public corpus (FERC release, CMU copy, via Hugging Face corbt/enron-emails): 140 at random from those mailboxes (250-2,200 characters, no newsletters) and 60 more that mention privilege, counsel or litigation, seed 30 and 31 200; 198 labelled, 2 left out as undecidable enron_labels.txt
Synthetic Harborline Freight: 28 emails written for this vertical (fictional people, .example domains) 28 synthetic.json, with expected flags, copies and near-duplicates
Planted leaks Privilege-log descriptions of the synthetic documents: 21 leaky (3 quote the document, 3 state its figures, 15 paraphrase the advice) and 16 clean 37 leak_planted.json

Who wrote the labels: Claude (an AI agent, Opus 5.5), on 25 Sep 2026, reading each message as shown (bodies cut at 1,400 characters), before any model run. No lawyer has reviewed them. The rules used are at the top of enron_labels.txt: privilege needs a lawyer acting as a lawyer, a request for or the giving of legal advice or legal services (drafting counts), and confidentiality; work product needs litigation, arbitration or an investigation behind it; footers, "privileged" in a subject and a copied lawyer decide nothing; anything to or from the other side is not privileged. Each label is marked clear or hard: 126 clear, 72 hard (judgment calls such as "a business person asks a lawyer to draft a form", where reasonable reviewers differ). Of the 198: 80 privileged (59 attorney-client, 21 both), 118 not. The public TREC 2010 Legal Track privilege judgments (topic 304) would be a stronger, expert-labelled set, but they are keyed to EDRM Enron v2 document ids that we could not map to this copy of the corpus in the time available.

Splits: seed 30 shuffles the 200 ids; the first 50 are dev (49 labelled), the other 150 are test (149 labelled). Dev was used to write and adjust the prompts and the review rules; test was run once, after the rules were frozen, and nothing was changed because of it. One change came from dev: the first version decided on the top option's probability, and dev showed that attorney-client and "both" split the mass on litigation emails (a document the model is sure is privileged scored 0.35 for its top option), so the decision moved to p(privileged) = 1 - p(not privileged), with the withhold and produce band at 0.7 / 0.3 chosen on dev. The synthetic set was written before any model run; two wording fixes to review reasons and one to a consistency message followed its first run, with no change to any call.

What is measured

  • Calls: the model's call against the label, as privileged-vs-not (what decides production) and four-way. p(privileged) is scored by AUROC (does it rank privileged above not privileged) and ECE (10 bins).
  • After the review rules: the documents the tool would produce although labelled privileged (the dangerous error: a waiver), the ones it would withhold although not privileged, and how many go to a lawyer.
  • Leaks: the leak check on the planted descriptions (code rules, judge, both), and how many drafted log descriptions leaked on the real runs.
  • Consistency: on the synthetic set, whether copies, near-duplicate drafts and threads got the same withhold decision where their labels agree, the expected flags and review calls, and a full re-run.

Results: direct route with logprobs (the self-host configuration)

One call per question, probabilities from the answer token's log-probabilities (typed-judgment calibration), about 5 calls per document: the privilege call, two element questions, the grounding check of the reason, and for withheld documents a description and the leak judge.

Set n (privileged) Model call: accuracy / precision / recall AUROC p(priv) ECE Privileged produced Wrongly withheld To review Auto-decided agree
Enron dev 49 (16) 95.9% / 93.8% / 93.8% 0.998 0.121 0 0 9 (18%) 40 of 40
Enron test (held out) 149 (64) 83.9% / 82.3% / 79.7% 0.908 0.058 4 (6.2% of privileged) 3 57 (38%) 85 of 92 (92.4%)
— test, clear labels 92 (30) 95.7% / 93.3% / 93.3% 0.987 0.109 0 0 24 (26%) 68 of 68
— test, hard labels 57 (34) 64.9% / 71.9% / 67.6% 0.670 0.091 4 3 33 (58%) 17 of 24
Synthetic 28 (19) 96.4% / 100% / 94.7% 1.000 0.121 0 0 3 25 of 25

Four-way agreement on the held-out set: 80.5% exact, Cohen's kappa 0.65 (0.83 on clear labels).

The misses. All four privileged emails it would have produced, and all three it would have withheld wrongly, are on labels we had marked hard:

  • produced: a lawyer's one-line agreement with an HR recommendation sent to someone whose role the email does not show; exposure figures a lawyer had asked for, forwarded to the general counsel; a business person asking a lawyer to draft an agency form; lawyer-drafted contract language passed around the business. The model read each as business content.
  • withheld: a lawyer asking who handles a charge (routing); a law firm's client alert on new CFTC rules (a general publication); an execution copy of an agreement sent by outside counsel. These are exactly the calls privilege reviewers argue about. The tool is not a substitute for that argument; it is a way to put the argument in front of a lawyer with the reason and the evidence.

Leaks on real drafts. 87 descriptions were drafted on the held-out set and 18 on the synthetic set. One first draft was flagged (it copied "online trading of forest products in Asia" from the email: a subject, arguably fine, rewritten anyway); none in the final log. The descriptions are deliberately generic ("Email from in-house counsel to a company employee providing legal advice regarding a commercial lease"). Planted leaks: the check caught 19 of 21 with 0 false alarms on 16 clean descriptions. The code rules alone caught 12 of 21 (all quotes and figures, 6 of 15 paraphrases); the judge caught 19. The two it missed are paraphrases that describe what was asked or approved ("asking in-house counsel to review a consultant's standard NDA before the company shares its lease figures"; "the CEO approves counsel's two proposed changes to the settlement").

Consistency (synthetic). The exact copy got the same call without a second model run; both near-duplicate pairs and all four threads got the same withhold decision where their labels agree (7 of 7 groups); all 12 expectations were met: the forward of privileged advice to an outside consultant and the advice sent to a personal account went to review with their waiver reasons, the CFO's summary of the advice to managers went to review because no lawyer is on it, the other side's settlement letter and the ops update with a lawyer only copied were produced. Re-run: 28 of 28 documents got the same withhold / review / produce decision; 27 of 28 the same privilege type (one reply moved from attorney-client to "both").

Cost and speed (direct route): 5.2 calls, about 3,850 prompt and 225 generated tokens per document; 1.4 s per document with 6 calls in flight on the shared card. At the gateway's list price for Qwen3.8-27B ($0.30 / $1.50 per million tokens) that is $1.49 per 1,000 documents; self-hosted it is GPU time.

Results: hosted gateway route (the hosted demo)

The gateway does not pass log-probabilities through yet, so the hosted route samples the privilege call (the temperature-0 answer plus 4 seeded samples, votes smoothed as in typed judgments) and asks the element and leak questions once each: about 10 calls per document.

Set n Model call: accuracy / precision / recall AUROC p(priv) ECE Privileged produced Wrongly withheld To review Expectations Calls / doc Cost per 1,000 docs
Synthetic 28 96.4% / 100% / 94.7% 0.974 0.060 0 0 3 12 of 12, 7 of 7 groups 8.9 $2.60

Same calls and the same review decisions as the direct route on this set. The held-out Enron set was not re-run on the hosted route: the gateway was shared with several other evaluation jobs during this session, and 28 documents took 1,420 s (about 50 s per document) at its worst, against 85-135 s for the 12-email demo when it was quiet. The latency on the Stack tab reports both. Hosted calls each carry a gateway-signed receipt (100 per 12-email demo run, $0.031 at list price).

Self-host check (25 Sep 2026): a fresh clone of the branch, docker build of docker/api/Dockerfile, the assemble prompt's api service (named volume) pointed at the already-running local vLLM with host networking: the image builds, the service is healthy, /privilege/info reports logprobs, the 12-email sample passes (17 s, 58 attested calls), the record verifies and fails when one call is changed, and the CSV export works. The model server's own startup was not re-verified.

Reading

  • On emails whose label is not a judgment call the tool agrees with the labels 96% of the time, and every one of its errors lands in the review band rather than in the produce pile. On the hard calls it is barely better than chance on its own (kappa 0.28), and the review rules catch most but not all of that: 4 of 34 hard privileged emails would have been produced. The review queue is the product. It cuts the documents a lawyer must read first by 60-75% on ordinary mail and puts the reasons, the lawyers and the evidence next to each one; it does not remove the lawyer.
  • p(privileged) ranks well (AUROC 0.91 held out) and is reasonably calibrated (ECE 0.06), so a firm can move the band to trade review load against risk on its own labelled sample.
  • The labels are one AI reviewer's reading, not a lawyer's, and the Enron sample is small. Publish your own measured precision and recall on a labelled sample of the matter before relying on it (and before quoting a per-document price).

Reproduce

python scripts/privilege_samples.py
python scripts/privilege_eval.py run --split dev --tag direct --base <direct-route server>
python scripts/privilege_eval.py run --split test --tag direct --base <direct-route server>
python scripts/privilege_eval.py run --split synthetic --tag direct [--repeat] --base <direct-route server>
python scripts/privilege_eval.py run --split synthetic --samples 4 --base <gateway-route server>
python scripts/privilege_eval.py leak --tag direct --base <direct-route server>
python scripts/privilege_eval.py metrics

The eval server needs DECOSA_PRIVILEGE_MAX_DOCS=50 for the 28-document synthetic set; the key is a dk_ key for privilege-log and typed-judgment in DECOSA_EVAL_KEY.