99 · Legal · live
Check their brief
Eval results
Scored on a held-out or test splitRun 28 Sep 2026Eval write-up (decosa-api, access required)
- Non-existent citations caught (strict)24 / 31 (77%)held outn = 31LePhantomCite eval split, run once on frozen code; 81% counting 'look at'.
- Case name and cite from two different cases (strict)34 / 68 (50%)held outn = 6875% counting 'look at'.
- Swapped-word misquotations (strict)31 / 45 (69%)held outn = 45
- Invented citations in real criticised filings called fine0 / 78test splitn = 7830 findings and 3 looks among the 34 it could look up; 32 not checked because CourtListener was unavailable (daily limit). Not held out from the fixes this run exposed.
- Misstated holdings (strict)4 / 131 (3%)held outn = 131Not reliably caught by the open judge. Our own citation-support model plus the judge: 80 / 129 at 7 / 372 false flags on held-out pairs (prototype, off by default).
- False findings on error-free excerpts49 / 950 (5.2%)held outn = 95033 / 950 (3.5%) with the post-test fixes simulated on the same outputs.
- Hidden AI instructions flagged (blind texts)71 / 74held outn = 74The 3 misses were CJK text the planting tool could not encode. Benign hidden texts flagged: 3 / 75.
- Cost per excerpt at list price$0.0123 USDheld outn = 390p50 9.9 s; the fictional two-page opposition: $0.0127, 32-37 s warm.
- Invented citations in criticised filings, with the local citation index (CourtListener off)32 / 78 findings, 63 / 78 findings or lookstest splitn = 780 called fine; 4 not checked (was 32 without the index). Not held out from the lookup fixes made during that run.
Dataset
LePhantomCite (CC BY 4.0): 390 held-out excerpts of real federal appellate briefs with injected citation errors; 160 blind-written injection and benign texts hidden in 12 public-domain Solicitor General briefs; 31 real filings courts criticised for invented citations and 23 uncriticised briefs from RECAP.
Caveats
- LePhantomCite errors are injected by its authors; its non-existent citations use reporter series that do not exist, which a code check catches.
- The published real-filings run had CourtListener from cache only (free-account daily limit); the re-run with the local citation index checks all but 97 of 2,030 case citations. Labels come from court orders, 17 of 31 of which list only examples.
- Pin cites and misstated holdings are not reliably caught by the shipped path.
- Some 'false findings' on error-free excerpts are real errors in the original briefs; not all were adjudicated.
- About 5% of test items met CourtListener rate limits (shared with our other jobs).
- One author wrote the rules and the dev texts; the test injection texts were blind-written.
- No human paralegal was timed; the manual baseline is an estimate.
Nightly smoke check
Loading the nightly status…
- Result
- pass
- Run
- 30 Sep 2026
- Latency, this run
- n/a
- p50 over passed runs
- 8.3 s
- Receipts
- 7
- Model calls
- n/a
- Tokens
- n/a
- Cost per run
- $0.013
Self-host verification
Verified on 28 Sep 2026: fresh clone, compose up, sample against local model servers
Fresh clone of a decosa-api pre-release build (6ee6b6b), the api image built from docker/api/Dockerfile (theirbrief extra), this prompt's compose with the direct route to the already-running local Qwen3.8-27B, OCR off, anonymous CourtListener. The prompt's smoke steps and the rehearsal bundle passed (11/11): 7 findings on the fictional PDF, 7 receipts, the record verifies and a tampered decision fails, the Word memo exports, a scan is refused with a clear message when OCR is off. 57.9 s, $0.0126. Torn down after.
Rehearsal bundle: check-their-brief.zip (4 KB, 11 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.
Known limits
- Hosted numbers are the production smoke check of the sample, run 5 times in a row on 30 Sep 2026 (all passed); with 5 runs the slowest-1-in-20 figure is simply the slowest run.
- No citator: it does not say whether a case is still good law.
- Westlaw- and Lexis-only decisions, many unpublished orders and most state codes are not in the free sources: they come back 'look at' or 'not checkable', never 'fine'.
- Case citations are looked up in a local index of CourtListener's and the Caselaw Access Project's public data (quarterly; snapshot 30 Jun 2026), with no network call. CourtListener's free API (250 searches a day) is asked only on a miss or for a volume newer than the snapshot; what nothing can answer is listed as 'not checked yet', never as fine.
- Quotations and holdings of cases after about 2018 need the opinion PDF from CourtListener, one search per quoted case; if it cannot be asked, those checks stay 'not checked'.
- Instructions hidden in the file are caught well; instructions written in plain view are caught only when a pattern or AI word flags the sentence first.
Models and licences
Standard tier. Licence posture: permissive (Apache, MIT or BSD).
- Checker: hidden-text scan (rendered page against text layer, metadata, comments, invisible Unicode), citation parsing, impossible-reporter check, lookups with a search trail, quotation match, Rule 5.2 scan, findings memo, signed record (no model; CPU)decosa-api check-their-brief (decosa_api/verticals/theirbrief, on the filing pre-flight engine)AGPL-3.0-or-later
- Judge: one call per holding checked against the opinion, and one call for which hidden or embedded texts speak to AI tools (texts quoted as data)Qwen3.8-27B (NVFP4)Apache-2.0
- Document reader for scanned filings (no text layer): layout plus OCR, then the same checksDecosa document reader (Docling layout + PaddleOCR-VL-1.6)Apache-2.0
All quality evidence
Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.
Lite · CPU only, no GPU (2)
- Hidden-text scan and reporter/existence checks: same as standard (code, no model)docs/evals/check-their-brief.md
- Hidden instructions without the model's reading: not measured separately
Standard · one GPU for the judge (hosted demo) (10)
- Citations to cases that do not exist, caught as a finding (strict) / as a finding or a look (lenient): 24/31 (77%) / 25/31 (81%)docs/evals/check-their-brief.md, LePhantomCite held-out test (390 real brief excerpts, CC BY 4.0), run once
- Case name and cite that belong to two different cases: 34/68 (50%) strict / 51/68 (75%) lenientsame
- Quotations with a word swapped: 31/45 (69%) strict / 36/45 (80%) lenientsame
- Wrong pin cites and misstated holdings (strict): 2/55 and 4/131: not reliably caughtsame; the judge marks most misstated holdings 'look at', as it does 25% of holdings in error-free excerpts
- False findings on error-free excerpts: 49 of 950 checked items (5.2%) as run; 33 (3.5%) with the post-test fixes simulated on the same outputssame
- Hidden instructions aimed at AI tools (blind-written texts hidden in 12 real briefs, 13 techniques): 71/74 flagged (the 3 misses were CJK text the planting tool could not encode); benign hidden texts flagged 3/75 as run, 1/75 in the regression re-run after the pattern fixdocs/evals/check-their-brief.md, section 3 (re-run after the real-PDF fixes: same numbers)
- Instructions written in plain view: 2/6 flagged, 0/5 benign flagged: weaksame
- Real filings courts criticised for invented citations (31 filings, 161 problems from the orders; CourtListener cache-only): Invented citations: 0 of 78 called fine; 30 findings and 3 looks among the 34 it could look up; 32 not checked because CourtListener was unavailable. All problems: 37/161 findings, 74/161 findings or looksdocs/evals/check-their-brief.md, section 2 (Charlotin database + RECAP; not held out from the fixes it exposed)
- Same 54 filings with the local citation index, CourtListener off (gateway, 28 Sep 2026): Invented citations: 0 of 78 called fine; 32 findings and 63 findings or looks; 4 not checked (was 32). Case citations not checked: 97 of 2,030 (was 563). Uncriticised briefs: 47 findings on 2,006 checked items (2.3%). p50 17 s a filing.docs/evals/citation-index.md (not held out from the lookup fixes made during that run)
- Findings on 23 uncriticised real briefs: 66 of 1,645 checked items (4.0%) as run; 49 of 1,666 (2.9%) after the fixes (direct-route re-run); many are real miscites in those briefs or quotes of the other side's invalid citesdocs/evals/check-their-brief.md, section 2
Wanted · a GLM-5.3-Flash holdings judge on your own hardware (1)
- This eval, same protocol: not measured yet