Skip to content
decosa

The full write-up behind the numbers on the tool’s page, as the team that built it wrote it: data, method, results and caveats. Internal names are removed; nothing else is edited.

Filing pre-flight: eval (25 Sep 2026)

Questions: (1) does the pre-flight catch fake citations, citations for a holding the case does not contain, and misquotations? (2) how often does it call something a "problem" in a real, careful brief?

Data

  • Real briefs: 12 briefs filed by the Office of the Solicitor General in 2025-2026 (justice.gov/osg; works of the US government, public domain), converted with pdftotext. For each, the first ~9,000 characters of the argument section (or "reasons for granting/denying") after cleaning. These are carefully cite-checked briefs, so we treat every "problem" on the unaltered text as a false problem unless we found it to be real (we found none that were real).
  • Dev (3 briefs: Case v. Montana amicus, Beck, Abdulla): used to fix the parser and set the statuses.
  • Test (9 briefs: Urias-Orellana, East Penn, Schena, CashCall, Asante, Save Jobs USA, Uriostegui-Hernandez, Ebu, FirstEnergy): run once, after the code was frozen, with scripts/preflight_eval.py --split test.
  • Planted errors (seed 17), in a copy of each excerpt:
    • fake_cite: the six fabricated authorities from Mata v. Avianca, Inc., 678 F. Supp. 3d 443 (S.D.N.Y. 2023) (test) or two invented ones (dev), each in its own sentence; plus two real full case cites per brief whose first page was moved by 7-40 pages.
    • wrong_prop: a sentence or parenthetical citing a well-known Supreme Court opinion for something it does not hold (15 hand-written for test, 5 different ones for dev; listed in the script). Most are reversals of the holding; some are subtler (Gates "reaffirming" Aguilar-Spinelli, Celotex requiring affidavits, Daubert applying Frye).
    • misquote: one word changed inside a quotation the clean run had verified (a meaning-changing swap such as may/must, not/also, is/was, or, when none is available, a singular/plural change).
  • A plant counts as caught when the item at its position is problem (strict) or problem/review (lenient).

Results: planted errors (test, held out)

Error Planted Problem Review Missed Strict Lenient
Fake citation 19 17 1 1 89% 95%
Citation for a holding the case does not contain 13 12 0 1 92% 92%
Misquotation (one word) 22 14 7 1 64% 95%
  • All six Mata v. Avianca fakes were caught, five as problems. Martinez v. Delta Airlines (a Westlaw cite) came back "review": no opinion carries the cite, but a docket of that name exists in CourtListener.
  • The missed fake: a moved page on a cert-denial cite with no case name ("cert. denied, 525 U.S. 1055"), which landed on another real order. Without a name there is nothing to cross-check.
  • The missed wrong proposition: the Gates parenthetical begins "reaffirming", a verb the parser does not treat as a holding, and the sentence left no claim to check.
  • The seven misquotes marked "review" are all singular/plural changes ("doffing" to "doffings"): the check labels a change of word form "spelling or word form differs" rather than "problem". Of the 15 meaning-changing swaps, 14 were problems and one was missed: the changed word ("is" to "was") sat inside the brief's own bracketed insertion ("[an applicant was]"), which the check skips by design, so the quoted words themselves were still verbatim.

Results: false problems on the unaltered briefs (test, 9 briefs, 511 checked items)

Check Items Problem Review Unverified OK
a. Authorities exist 110 7 (6.4%) 14 14 74
b. Propositions 111 1 (0.9%) 52 (47%) 24 29
c. Record cites 92 0 0 87 0
d. Quotations 198 11 (5.6%) 1 123 63
All 511 19 (3.7%) 67 (13%) 248 166

What the 19 false problems were:

  • Authorities (7): three Supreme Court orders from 2021-2024 cited by name at their S. Ct. page, which the free indexes do not hold; one 2025 Westlaw cite to a decision not in CourtListener under that cite or name, counted three times because the brief cites it three times; one party name cut short by PDF line breaks ("Comm' v. FERC"), which then failed the name match.
  • Quotations (11): every one was a quotation attributed to the wrong source: words quoted from the record (a trial transcript), from a statute, or from another case, sitting next to a case citation. The words were not in that case, which is true, but the brief never said they were.
  • Propositions (1): the judge read Anderson v. Mt. Clemens Pottery's de minimis discussion as contradicting a sentence about what the Court held; arguably the brief's sentence is a fair reading.
  • Record cites were "unverified" because the eval gave no record excerpts (the brief cites the petition appendix); record resolution is covered by unit tests and the fictional demo.

"Unverified" is large by design: quotations from the record, statutes the free sources do not hold, law reviews, books, and citations that could not be tied to a source are never counted as OK.

Propositions are a triage list, not a verdict. Of the 82 propositions the judge read on the test briefs, it found 29 supported, 24 partly supported, 25 not supported by the passages read, and 4 contradicted. Most "not supported" verdicts on these real briefs are sentences that argue from a case rather than report what it held; they stay "review" unless the sentence attributes a holding to the court ("the Court held that ...", "(holding that ...)"), which is also exactly where the planted wrong propositions sit.

Dev split (for reference; the statuses were set on it)

Planted: fake citations 7/8 problem (1 review), wrong propositions 5/5, misquotes 7/9 (2 review). False problems on the unaltered dev briefs: 1 of 240 items (a quotation from a Supreme Court slip-opinion PDF where a footnote splits the passage; PDF-only misses are now "review"). Two changes were made after the dev run and before the test run (a case found by name at another page of the same volume is a wrong cite; quotations missing from PDF text are "review"), and two after the test run that do not change any status (masking identifiers in exports; retrying CourtListener 429s).

Privacy check

  • Patterns (SSN, taxpayer number, birth date next to "born"/"DOB", account numbers next to "account"/"card", Luhn-valid card numbers): 0 findings on the full text of all 12 real briefs (570,911 characters). Names next to child cues produced 5 candidates ("Immigration Appeals", "El Salvador", ...), which go to the model to confirm; none is a person.
  • The fictional demo brief's four planted identifiers (SSN, birth date, account number, a 14-year-old's full name) are all found, and the unit tests cover masking in every export and in the signed record.

Latency (our server, shared RTX PRO 6000, gateway route)

A 9,000-character argument section took 12-157 s (median about 35 s on test), almost all of it in judge calls and CourtListener lookups; the fictional demo takes 40-95 s depending on gateway load. Anonymous CourtListener use hit HTTP 429 during the eval; those lookups become "unverified". A free CourtListener account token (DECOSA_PREFLIGHT_COURTLISTENER_TOKEN) raises the limit.

Limits of this eval

  • Small: 12 briefs, 54 planted errors on test. The rates have wide intervals (Wilson 95%: 17/19 fake citations 0.69-0.97; 12/13 wrong propositions 0.67-0.99; 19/511 false problems 0.02-0.06).
  • One author wrote the wrong propositions, and they lean toward clear reversals; subtle mischaracterisations are harder.
  • The real briefs are from one careful filer. Briefs from smaller firms cite more unpublished and Westlaw-only decisions, which the free sources cover less well.
  • No human cite-checker was timed against it.

Verdict

Would a buyer pay? For the bundle, yes, as a pre-filing triage step: it reliably catches invented cases and changed quotations, points at the exact words, and adds the record-cite and Rule 5.2 checks nobody else bundles, with a signed record for the file, on hardware a firm can own. What is missing before it is a product rather than a demo: a good-law (citator) signal, which free sources do not give; better coverage of recent and Westlaw-only decisions (a CourtListener token helps; a paid database would need a licence); a cleaner quotation-to-source attribution for quotes of the record and statutes (the main source of false problems); the judge's standing-order certification text; and a lighter "review" load on propositions (a stronger judge, or showing only sentences that attribute a holding).

With the local citation index (28 Sep 2026)

Test split re-run on the same excerpts and planted texts with the local case-law citation index on and CourtListener off (gateway): planted fake citations 17 problems + 1 look of 19 (as with CourtListener live); unaltered briefs 1 authority problem (was 7), 15 unverified (14 live; 40 with neither). p50 13.5 s per excerpt. See docs/evals/citation-index.md.