Skip to content
decosa

30 · Legal · live

Privilege review and privilege log

Open the toolJSON

Eval results

Scored on a held-out or test splitRun 25 Sep 2026Eval write-up (decosa-api, access required)

  • Privileged vs not, model call: accuracy / precision / recall83.9% / 82.3% / 79.7%test splitn = 149Direct route with logprobs. Clear labels (92): 95.7%; hard labels (57): 64.9%. Dev: 95.9%.
  • AUROC of p(privileged) / ECE0.908 / 0.058test splitn = 149
  • Privileged emails the tool would have produced (waiver risk)4 (6.2% of privileged)test splitn = 64All four on labels marked hard. Wrongly withheld: 3.
  • Sent to attorney review57 (38%)test splitn = 149Auto-decided documents agreeing with the labels: 85 of 92 (92.4%).
  • Four-way call: exact / Cohen's kappa80.5% / 0.65test splitn = 149Kappa 0.83 on clear labels
  • Planted leaky log descriptions caught19 of 21syntheticn = 370 false alarms on 16 clean descriptions; code rules alone 12 of 21.

Dataset

200 emails from Enron lawyers' mailboxes in the public corpus (198 labelled; 49 dev, 149 test), a 28-email synthetic set and 37 planted log descriptions. Dev was used to write the prompts and review rules; test was run once after the rules were frozen, with nothing changed because of it.

Caveats

  • Labels are one AI reviewer's (Claude), not a lawyer's; 72 of 198 are marked hard judgment calls.
  • Small sample; publish your own measured precision and recall on a labelled sample of the matter before relying on it.
  • The held-out Enron set was run on the direct (self-host, logprobs) route only; the hosted gateway route was measured only on the 28 synthetic emails.
  • On hard calls the model alone is barely better than chance (kappa 0.28); the review queue, not the model, is the safeguard.

Nightly smoke check

Loading the nightly status…

Result
pass
Run
25 Sep 2026
Latency, this run
n/a
p50 over passed runs
121 s
Receipts
100
Model calls
n/a
Tokens
n/a
Cost per run
$0.031

Self-host verification

Verified on 25 Sep 2026: Fresh clone of the branch into a clean directory, docker build of docker/api/Dockerfile, the assemble prompt's api service with a named volume, pointed at the already-running local vLLM (Qwen3.8-27B NVFP4 on 127.0.0.1:8114) through host networking; then torn down.

Verified on 2026-09-25: the image builds, the service starts healthy, info reports logprobs, the harborline-dispute sample passes end to end (copy of HL-002 gets its call, HL-003 goes to review, HL-012 and HL-021 produced, 17 s, 58 attested calls), the signed record verifies and fails when one call is changed, and the CSV export works. The model server's own startup was not re-verified (no new GPU load).

Rehearsal bundle: privilege-log.zip (3 KB, 11 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.

Known limits

  • Labels are one AI reviewer's (Claude's), not a lawyer's; hard judgment calls are where it errs: 4 of 34 hard privileged held-out emails would have been produced.
  • About 38% of held-out real email goes to attorney review (26% where the label is clear).
  • The hosted gateway route slows sharply when the shared gateway is loaded (one 12-email run took 938 s).
  • Attachments must be sent as separate documents; no OCR, PDF or native file parsing in this version.
  • Partial privilege (redacting part of a document) is not proposed; such documents go to review.

Models and licences

Standard tier. Licence posture: permissive (Apache, MIT or BSD).

  • Reviewer: people map, waiver flags, duplicate and thread grouping, review rules, leak rules, consistency, signed record and ledger (no model; CPU)decosa-api privilege module (decosa_api/verticals/privilege)AGPL-3.0-or-later
  • Model: the typed privilege call, the two element questions, the grounding check of the reason, the log description and the leak judgeQwen3.8-27B (NVFP4)Apache-2.0

All quality evidence

Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.

Lite · one 32 GB card, self-hosted (1)
  • Calls on held-out Enron email: not measured separately: the same weights and prompts as the standard tier, so the calls should match; speed on a 5090 not measurednot measured yet
Standard · one 96 GB card (measured; hosted demo) (8)
  • Privileged vs not, model call, 149 held-out Enron emails (64 privileged by our labels): accuracy / precision / recall / AUROC of p(privileged): 83.9% / 82.3% / 79.7% / 0.908docs/evals/privilege-log.md, direct route with logprobs; labels written by Claude (an AI agent), not a lawyer
  • After the review rules: privileged emails the tool would have produced / non-privileged it would have withheld / sent to attorney review: 4 of 64 (6.2%) / 3 of 85 / 57 of 149 (38%); decided without review: 92, of which 92.4% agree with the labelsdocs/evals/privilege-log.md, same run
  • The same, on the 92 held-out emails whose label we marked clear (not a judgment call): accuracy 95.7%, 0 privileged produced, 0 wrongly withheld, 26% to reviewdocs/evals/privilege-log.md; all 4 misses and 3 over-withholds were on the 57 emails we marked as hard calls
  • Four-way call (attorney-client / work product / both / not privileged), held out: 80.5% exact, Cohen's kappa 0.65docs/evals/privilege-log.md
  • Planted leaky log descriptions caught (21 leaky, 16 clean, synthetic documents): 19 of 21 caught, 0 of 16 false alarms (code rules alone 12 of 21; the judge catches paraphrases)docs/evals/privilege-log.md, leak-check stage, direct route
  • Drafted log descriptions that leaked (87 on held-out Enron, 18 on the synthetic set): 1 first draft flagged (a subject phrase copied from the email), rewritten; 0 in the final log by the code rulesdocs/evals/privilege-log.md
  • Synthetic set (28 emails): labels met / expected flags and groups met / consistency across copies, drafts and threads / same decision on a re-run: 0 privileged produced, 0 wrongly withheld, 3 to review / 12 of 12 / 7 of 7 groups consistent / 28 of 28 (27 of 28 same privilege type)docs/evals/privilege-log.md
  • Hosted gateway route on the synthetic set (sampled call): privileged produced / wrongly withheld / to review / expectations met: 0 / 0 / 3 / 12 of 12; same privileged-vs-not accuracy as the direct route (96.4%), AUROC 0.974docs/evals/privilege-log.md, hosted route stage
Wanted · a larger second judge on your own hardware (1)
  • accuracy and AUROC, same protocol as standard: not measured yet

How we measure · All tools