21 · Legal · live
Filing pre-flight
Eval results
Scored on a held-out or test splitRun 25 Sep 2026Eval write-up (decosa-api, access required)
- Fake citations caught as problem (strict)89%test splitn = 1917 of 19; 95% lenient (problem or review). Wilson 95% 0.69-0.97.
- Citations for a holding the case does not contain, caught92%test splitn = 1312 of 13, strict and lenient.
- One-word misquotations caught as problem (strict)64%test splitn = 2295% lenient; the 7 'review' ones are singular/plural changes.
- False problems on 9 unaltered real briefs19 of 511 (3.7%)test splitn = 511Plus 67 items (13%) marked review; Wilson 95% 0.02-0.06.
- Privacy pattern findings on 12 real briefs0test splitn = 12570,911 characters of public-domain briefs.
Dataset
12 real Solicitor General briefs from 2025-2026 (public domain; 3 dev, 9 test, first ~9,000 characters of the argument) with planted errors: Mata v. Avianca fake cites, moved pages, hand-written wrong propositions and one-word misquotes.
Caveats
- Small: 12 briefs and 54 planted errors on test, so the rates have wide intervals.
- One author wrote the wrong propositions, and they lean toward clear reversals; subtle mischaracterisations are harder.
- The real briefs come from one careful filer; briefs citing more unpublished or Westlaw-only decisions are covered less well.
- Propositions are a triage list: 47% of real-brief propositions came back 'review'.
- No human cite-checker was timed against it, and there is no good-law (citator) signal.
Nightly smoke check
Loading the nightly status…
- Result
- pass
- Run
- 29 Sep 2026
- Latency, this run
- n/a
- p50 over passed runs
- 12 s
- Receipts
- 15
- Model calls
- n/a
- Tokens
- n/a
- Cost per run
- $0.015
Self-host verification
Verified on 25 Sep 2026: fresh clone, compose up, sample against local model servers
Method: a fresh clone of decosa-api main, the api image built from it, the compose file from this prompt, then the prompt's smoke steps and the nightly smoke module, against the already-running local Qwen3.8-27B vLLM. Verified on 25 Sep 2026: images build, services start, sample passes end to end against local model servers equivalent to the documented ones; model-server startup itself not re-verified. Offline (the default in the prompt) the planted record-cite and privacy problems are found and the cases come back unverified, as documented; with lookups on, the same 9 problems as hosted. The signed record verifies and a tampered entry is named.
Rehearsal bundle: filing-preflight.zip (6 KB, 12 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.
Known limits
- No citator: it does not say whether a case is still good law. Westlaw- or Lexis-only decisions, many unpublished orders, agency decisions and state codes come back unverified, never OK.
- Hosted speed depends on load on the shared service: the time shown is the median of our latest production runs of the fictional sample.
Models and licences
Standard tier. Licence posture: permissive (Apache, MIT or BSD).
- Checker: citation parsing, lookups, quotation matching, record cites, privacy scan, signed record (no model; CPU)decosa-api filing pre-flight (decosa_api/verticals/preflight)AGPL-3.0-or-later
- Judge: one call per proposition and per record cite, plus one call for minors' namesQwen3.8-27B (NVFP4)Apache-2.0
All quality evidence
Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.
Lite · CPU only, no GPU (3)
- Fake citations caught (existence and name checks need no model): 17/19 problem, same as standarddocs/evals/filing-preflight.md, 9 held-out Solicitor General briefs; the existence check is identical without the judge
- One-word misquotations (string match, no model): 14/22 problem, 21/22 problem or review, same as standarddocs/evals/filing-preflight.md, 9 held-out Solicitor General briefs
- Citations for a holding the case does not contain: not checked in this tierno judge
Standard · one GPU for the judge (hosted demo) (6)
- Fake citations caught (the six Mata v. Avianca fakes plus real cites with the first page moved): problem / problem or review: 17/19 (89%) / 18/19 (95%)docs/evals/filing-preflight.md, 9 held-out Solicitor General briefs
- Citations for a holding the case does not contain, caught as problem: 12/13 (92%)docs/evals/filing-preflight.md, 9 held-out Solicitor General briefs
- One-word misquotations: problem / problem or review: 14/22 (64%) / 21/22 (95%); the 7 reviews are singular/plural changesdocs/evals/filing-preflight.md, 9 held-out Solicitor General briefs
- False problems on the same briefs unaltered (511 checked items): 19 (3.7%): quotations tied to the wrong source 11, authorities the free indexes lack 6, name parse 1, proposition 1docs/evals/filing-preflight.md, 9 held-out Solicitor General briefs
- Propositions read by the judge on real briefs: supported / partly / not supported by passages read / contradicted: 29 / 24 / 25 / 4 of 82 (all but 1 shown as review, not problem)docs/evals/filing-preflight.md, 9 held-out Solicitor General briefs
- Privacy patterns on the full text of 12 real briefs (570,911 characters): 0 false findingsdocs/evals/filing-preflight.md
Wanted · a GLM-5.3-Flash judge on your own hardware (1)
- This eval, same protocol: not measured yet