34 · Public sector · Legal · live
Public-records desk
Eval results
Scored on a held-out or test splitRun 25 Sep 2026Eval write-up (decosa-api, access required)
- Personal-data spans redacted (PII recall), final pipeline59 of 60 (98.3%)test splitn = 60Test run 5; the test sets were run five times with fixes in between.
- Planted spans redacted, final pipeline70 of 71 (98.6%)test splitn = 71Pattern finders alone: 33 of 71 (46.5%).
- Redaction precision / keep violations83.1% (69 of 83) / 3test splitn = 83Most over-redactions are company officials' names and business emails.
- Responsiveness precision / recall96.2% / 100.0%test splitn = 77
- Withheld spans recoverable from the release PDF0 of 95test splitn = 95All readers; an earlier run (test run 2) failed this check.
- PII recall on a set new to the pipeline (bus-contract, run 3)5/8held outn = 8Responsiveness 7/8; the cleanest held-out number, before the medical-span fix.
- Cost at list price$0.77 per 1,000 recordstest splitn = 77
Dataset
Synthetic agency records (3 dev sets, 4 held-out sets, 48 held-out records) labelled by Claude, plus 29 public FERC-released Enron emails labelled by hand for responsiveness.
Caveats
- The test sets were run four times and three design faults were fixed after seeing test results; the final numbers are not clean held-out numbers.
- Labels were written by one AI reviewer (Claude), not a records officer.
- Small sample: 48 synthetic held-out records and 29 public emails.
- Synthetic records are cleaner than real ones: no OCR noise, attachments or long threads.
- Medical details are the weakest kind (5/7); an officer still has to read each released record.
- Only the hosted route was measured; the self-host route was smoke-tested.
Nightly smoke check
Loading the nightly status…
- Result
- pass
- Run
- 25 Sep 2026
- Latency, this run
- n/a
- p50 over passed runs
- 40 s
- Receipts
- 41
- Model calls
- n/a
- Tokens
- n/a
- Cost per run
- $0.012
Self-host verification
Verified on 25 Sep 2026: Fresh clone of the branch into a clean directory, docker build of docker/api/Dockerfile, the assemble prompt's api service with a named volume, pointed at the already-running local vLLM (Qwen3.8-27B NVFP4 on 127.0.0.1:8114) through host networking; then torn down.
Verified on 2026-09-25: the image builds, the service starts healthy, info reports logprobs and both exemption lists, scoping works, the coastal-permits sample passes end to end (17 s, 41 calls, every disposition as expected, release check ok on 28 spans), the signed record verifies and fails at the changed entry, the PDF export carries no withheld text, one officer decision rebuilds the release and the ledger verifies. The model server's own startup was not re-verified (no new GPU load).
Rehearsal bundle: foia-desk.zip (4 KB, 11 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.
Known limits
- Labels are one AI reviewer's (Claude's), not a records officer's; the held-out sets are small and synthetic.
- Medical details are the weakest kind: both held-out misses were medical.
- It over-redacts company officials' names that the model reads as private, and agency staff named only in a record's text (not in its mail headers); they go to review or to the officer.
- Runs are not identical: the same records can come back with a few redactions more or fewer from one run to the next (temperature 0 on a shared server is not bit-for-bit repeatable). The pattern finders, the date filter and the last contact check are code and do not vary; the officer decides every proposal.
- The release is re-typeset from text: no original layout, no PDF, scan or email-archive intake, no audio or video.
- The hosted route slows when the shared gateway is loaded.
Models and licences
Standard tier. Licence posture: permissive (Apache, MIT or BSD).
- Records desk: date filter, pattern finders, exemption lists and reasons, reason leak check, consistency, release PDF writer and its integrity check, index, letter, signed record and ledger (no model; CPU)decosa-api foia module (decosa_api/verticals/foia)AGPL-3.0-or-later
- Model: the request scope, the responsiveness and deliberative calls, the redaction spans with their category, description and harm, and the privilege engine's callsQwen3.8-27B (NVFP4)Apache-2.0
All quality evidence
Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.
Lite · one 32 GB card, self-hosted (1)
- Held-out test sets: not measured separately: the same weights and prompts as the standard tier; speed on a 5090 not measurednot measured yet
Standard · one 96 GB card (measured; hosted demo) (7)
- Planted personal details redacted, 48 held-out synthetic records (60 labelled: names, phones, addresses, emails, SSN, dates of birth, licence and account numbers, medical details): 59 of 60 (98.3%); the pattern finders alone would catch 46.5% of all planted spansdocs/evals/foia-desk.md, test run 5 (30 Sep 2026, current pipeline); labels written by Claude (an AI agent), not a records officer
- What it missed: The misses are listed per record in the eval write-up; medical details are the weakest kind.docs/evals/foia-desk.md
- Responsiveness, 77 held-out records (48 synthetic, 29 public FERC-released Enron emails): precision / recall: 96.2% / 100.0%docs/evals/foia-desk.md, test run 5 (30 Sep 2026, current pipeline)
- Exemption label on redacted spans / records with agency counsel withheld in full with the right exemption: 70 of 70 / 6 of 6 (federal (b)(5), (b)(6), (b)(4); CPRA § 7927.700, § 7927.705, § 7922.000)docs/evals/foia-desk.md, test run 5 (30 Sep 2026, current pipeline)
- Redaction precision (redactions that overlap a labelled span) / text that had to be released but was boxed: 83.1% / 3 spans: mostly company officials' names and emails, boxed everywhere by the privacy-protective consistency step and sent to reviewdocs/evals/foia-desk.md, test run 5 (30 Sep 2026, current pipeline)
- Release check: withheld spans recoverable from the PDF (built-in reader, pdfplumber, filing-preflight black-box check): 0 of 95 on the held-out releases; 0 characters under any boxdocs/evals/foia-desk.md, test run 5 (30 Sep 2026, current pipeline)
- How the numbers moved as design faults were fixed (the test sets were run five times; no threshold tuned on them): personal details redacted 81.0% → 78.8% → 91.7% → 96.7% → 98.3%; responsiveness recall 91.4% → 86.0% → 98.0% → 100.0% → 100.0%docs/evals/foia-desk.md, History; two fresh held-out sets were written before runs 2 and 3