Skip to content
decosa

99 · Legal · live

Check their brief

Open the toolJSON

Eval results

Scored on a held-out or test splitRun 28 Sep 2026Eval write-up (decosa-api, access required)

  • Non-existent citations caught (strict)24 / 31 (77%)held outn = 31LePhantomCite eval split, run once on frozen code; 81% counting 'look at'.
  • Case name and cite from two different cases (strict)34 / 68 (50%)held outn = 6875% counting 'look at'.
  • Swapped-word misquotations (strict)31 / 45 (69%)held outn = 45
  • Invented citations in real criticised filings called fine0 / 78test splitn = 7830 findings and 3 looks among the 34 it could look up; 32 not checked because CourtListener was unavailable (daily limit). Not held out from the fixes this run exposed.
  • Misstated holdings (strict)4 / 131 (3%)held outn = 131Not reliably caught by the open judge. Our own citation-support model plus the judge: 80 / 129 at 7 / 372 false flags on held-out pairs (prototype, off by default).
  • False findings on error-free excerpts49 / 950 (5.2%)held outn = 95033 / 950 (3.5%) with the post-test fixes simulated on the same outputs.
  • Hidden AI instructions flagged (blind texts)71 / 74held outn = 74The 3 misses were CJK text the planting tool could not encode. Benign hidden texts flagged: 3 / 75.
  • Cost per excerpt at list price$0.0123 USDheld outn = 390p50 9.9 s; the fictional two-page opposition: $0.0127, 32-37 s warm.
  • Invented citations in criticised filings, with the local citation index (CourtListener off)32 / 78 findings, 63 / 78 findings or lookstest splitn = 780 called fine; 4 not checked (was 32 without the index). Not held out from the lookup fixes made during that run.

Dataset

LePhantomCite (CC BY 4.0): 390 held-out excerpts of real federal appellate briefs with injected citation errors; 160 blind-written injection and benign texts hidden in 12 public-domain Solicitor General briefs; 31 real filings courts criticised for invented citations and 23 uncriticised briefs from RECAP.

Caveats

  • LePhantomCite errors are injected by its authors; its non-existent citations use reporter series that do not exist, which a code check catches.
  • The published real-filings run had CourtListener from cache only (free-account daily limit); the re-run with the local citation index checks all but 97 of 2,030 case citations. Labels come from court orders, 17 of 31 of which list only examples.
  • Pin cites and misstated holdings are not reliably caught by the shipped path.
  • Some 'false findings' on error-free excerpts are real errors in the original briefs; not all were adjudicated.
  • About 5% of test items met CourtListener rate limits (shared with our other jobs).
  • One author wrote the rules and the dev texts; the test injection texts were blind-written.
  • No human paralegal was timed; the manual baseline is an estimate.

Nightly smoke check

Loading the nightly status…

Result
pass
Run
30 Sep 2026
Latency, this run
n/a
p50 over passed runs
8.3 s
Receipts
7
Model calls
n/a
Tokens
n/a
Cost per run
$0.013

Self-host verification

Verified on 28 Sep 2026: fresh clone, compose up, sample against local model servers

Fresh clone of a decosa-api pre-release build (6ee6b6b), the api image built from docker/api/Dockerfile (theirbrief extra), this prompt's compose with the direct route to the already-running local Qwen3.8-27B, OCR off, anonymous CourtListener. The prompt's smoke steps and the rehearsal bundle passed (11/11): 7 findings on the fictional PDF, 7 receipts, the record verifies and a tampered decision fails, the Word memo exports, a scan is refused with a clear message when OCR is off. 57.9 s, $0.0126. Torn down after.

Rehearsal bundle: check-their-brief.zip (4 KB, 11 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.

Known limits

  • Hosted numbers are the production smoke check of the sample, run 5 times in a row on 30 Sep 2026 (all passed); with 5 runs the slowest-1-in-20 figure is simply the slowest run.
  • No citator: it does not say whether a case is still good law.
  • Westlaw- and Lexis-only decisions, many unpublished orders and most state codes are not in the free sources: they come back 'look at' or 'not checkable', never 'fine'.
  • Case citations are looked up in a local index of CourtListener's and the Caselaw Access Project's public data (quarterly; snapshot 30 Jun 2026), with no network call. CourtListener's free API (250 searches a day) is asked only on a miss or for a volume newer than the snapshot; what nothing can answer is listed as 'not checked yet', never as fine.
  • Quotations and holdings of cases after about 2018 need the opinion PDF from CourtListener, one search per quoted case; if it cannot be asked, those checks stay 'not checked'.
  • Instructions hidden in the file are caught well; instructions written in plain view are caught only when a pattern or AI word flags the sentence first.

Models and licences

Standard tier. Licence posture: permissive (Apache, MIT or BSD).

  • Checker: hidden-text scan (rendered page against text layer, metadata, comments, invisible Unicode), citation parsing, impossible-reporter check, lookups with a search trail, quotation match, Rule 5.2 scan, findings memo, signed record (no model; CPU)decosa-api check-their-brief (decosa_api/verticals/theirbrief, on the filing pre-flight engine)AGPL-3.0-or-later
  • Judge: one call per holding checked against the opinion, and one call for which hidden or embedded texts speak to AI tools (texts quoted as data)Qwen3.8-27B (NVFP4)Apache-2.0
  • Document reader for scanned filings (no text layer): layout plus OCR, then the same checksDecosa document reader (Docling layout + PaddleOCR-VL-1.6)Apache-2.0

All quality evidence

Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.

Lite · CPU only, no GPU (2)
  • Hidden-text scan and reporter/existence checks: same as standard (code, no model)docs/evals/check-their-brief.md
  • Hidden instructions without the model's reading: not measured separately
Standard · one GPU for the judge (hosted demo) (10)
  • Citations to cases that do not exist, caught as a finding (strict) / as a finding or a look (lenient): 24/31 (77%) / 25/31 (81%)docs/evals/check-their-brief.md, LePhantomCite held-out test (390 real brief excerpts, CC BY 4.0), run once
  • Case name and cite that belong to two different cases: 34/68 (50%) strict / 51/68 (75%) lenientsame
  • Quotations with a word swapped: 31/45 (69%) strict / 36/45 (80%) lenientsame
  • Wrong pin cites and misstated holdings (strict): 2/55 and 4/131: not reliably caughtsame; the judge marks most misstated holdings 'look at', as it does 25% of holdings in error-free excerpts
  • False findings on error-free excerpts: 49 of 950 checked items (5.2%) as run; 33 (3.5%) with the post-test fixes simulated on the same outputssame
  • Hidden instructions aimed at AI tools (blind-written texts hidden in 12 real briefs, 13 techniques): 71/74 flagged (the 3 misses were CJK text the planting tool could not encode); benign hidden texts flagged 3/75 as run, 1/75 in the regression re-run after the pattern fixdocs/evals/check-their-brief.md, section 3 (re-run after the real-PDF fixes: same numbers)
  • Instructions written in plain view: 2/6 flagged, 0/5 benign flagged: weaksame
  • Real filings courts criticised for invented citations (31 filings, 161 problems from the orders; CourtListener cache-only): Invented citations: 0 of 78 called fine; 30 findings and 3 looks among the 34 it could look up; 32 not checked because CourtListener was unavailable. All problems: 37/161 findings, 74/161 findings or looksdocs/evals/check-their-brief.md, section 2 (Charlotin database + RECAP; not held out from the fixes it exposed)
  • Same 54 filings with the local citation index, CourtListener off (gateway, 28 Sep 2026): Invented citations: 0 of 78 called fine; 32 findings and 63 findings or looks; 4 not checked (was 32). Case citations not checked: 97 of 2,030 (was 563). Uncriticised briefs: 47 findings on 2,006 checked items (2.3%). p50 17 s a filing.docs/evals/citation-index.md (not held out from the lookup fixes made during that run)
  • Findings on 23 uncriticised real briefs: 66 of 1,645 checked items (4.0%) as run; 49 of 1,666 (2.9%) after the fixes (direct-route re-run); many are real miscites in those briefs or quotes of the other side's invalid citesdocs/evals/check-their-brief.md, section 2
Wanted · a GLM-5.3-Flash holdings judge on your own hardware (1)
  • This eval, same protocol: not measured yet

How we measure · All tools