Skip to content
decosa

The full write-up behind the numbers on the tool’s page, as the team that built it wrote it: data, method, results and caveats. Internal names are removed; nothing else is edited.

Eval: public-records desk (use case 34)

Run on 25 Sep 2026 on our server, hosted gateway route (Qwen3.8-27B NVFP4 through the model gateway, one call per typed question with the stated confidence), on a GPU and gateway shared with other workloads' work. Scripts: scripts/foia_samples.py (development sets), scripts/foia_testsets.py (held-out sets), scripts/foia_eval.py (runs and scores). Score files: docs/evals/foia-desk/*-scores.json; the raw event logs are on our server in ~/data/foia-eval/runs/.

Data

Set Kind Records Planted spans Notes
coastal-permits dev (console sample), federal 12 22 resident complaints, SSNs, a deliberative email, agency counsel, a copy, out-of-range and off-topic records
alder-point dev (console sample), CA CPRA 9 14 public comments with medical details, a draft memo, the city attorney
ferc-enron dev (console sample), public 10 – FERC-released Enron email, responsiveness only
broadband-grants held out, federal 14 26 grants office; bank details, a sign-in sheet, an SSN form, a revised near-duplicate memo
water-rates held out, CA CPRA 14 24 water district; shutoff list, medical details, a copy of a draft memo
park-concessions held out, federal 10 12 written after test run 1 (see History)
bus-contract held out, CA CPRA 10 9 written after test run 2 (see History)
ferc-enron-test held out, public 29 – responsiveness only, 11 responsive

All synthetic names, numbers and domains are fictional (.example domains, 555-01xx phones). Labels (responsive or not, each span that should be redacted with its category, records to withhold in full, and "keep" text that must be released: official contacts, facts in deliberative records) were written by Claude (an AI agent) with the records, not by a records officer. The Enron emails were labelled by hand for a hypothetical request ("price caps or price mitigation in the California or western electricity markets, May to July 2001"); 9 ambiguous emails were dropped before labelling.

Metrics

  • Responsiveness: a record counts as passed on when its disposition is anything but "not responsive"; precision and recall against the labels.
  • Planted recall: a labelled span is redacted when redactions cover at least 80% of its characters, at every place it occurs in the record. "PII recall" is the same without the deliberative spans. Records labelled "withhold in full" are scored on the disposition and exemption instead.
  • Exemption labels: for each redacted span, whether the exemption on the covering redaction is the one the label's category maps to in the set's list; for withheld records, whether the cited exemption matches.
  • Redaction precision: the share of redactions in released records that overlap a labelled span. "Keep violations" are redactions over text labelled must-release.
  • Integrity: the release PDF is read back by the built-in reader (every drawn string and every literal in the file), pdfplumber's text layer and the filing-preflight black-box check; a withheld span is "recovered" when its words (or, for numbers, its digits) appear in any reader's text.
  • Pattern baseline: recall of the pattern finders alone, no model.

Final results (pipeline as shipped; test run 4)

Held-out sets (77 records) Dev sets (31 records)
Responsiveness precision / recall 98.1% / 100% (51 TP, 1 FP, 0 FN, 25 TN) 95.5% / 100%
Planted spans redacted 69 of 71 (97.2%) 36 of 36
Personal-data spans redacted (PII recall) 58 of 60 (96.7%) 30 of 30
Pattern finders alone 33 of 71 (46.5%) 17 of 36
Deliberative spans redacted 11 of 11 6 of 6
Exemption label correct on redacted spans 69 of 69 36 of 36
Records with counsel withheld in full, right exemption 6 of 6 4 of 4
Redaction precision / keep violations 78.9% (71 of 90) / 3 94.7% / 0
Release check: withheld spans recoverable 0 of 102 (all readers; 0 characters under boxes) 0 of 46
Model calls / cost at list price 216 calls, $0.059: $0.76 per 1,000 records $0.94 per 1,000
Wall time (shared gateway) 178 s for 77 records 157 s for 31

By kind (held out): names 19/19, phones 9/9, addresses 7/7, emails 6/6, account numbers 5/5, dates of birth 4/4, driver's licences 2/2, SSN 1/1, medical 5/7.

What it misses. Both misses are medical details: (1) WR-12, "our meter reader fractured his wrist on the job ... off on workers' compensation for six weeks": the worker is unnamed, and the model released the record in full, although coworkers could identify him; (2) WR-01, "My mother lives with me and uses an oxygen concentrator around the clock": 61% covered, the clause saying who lives there was left. Medical details are the weakest kind in every run (the model tends to box the diagnosis, not the words around it), so the console flags nothing extra for them today: an officer's read of each released record is still required.

Over-redaction. Most unlabelled redactions are company officials' names and business emails (a co-op's general manager, a concessionaire's manager, a bus company's contracts manager): the model called the person private in one record, and the consistency step, which is privacy-protective by design, then boxed the same words everywhere and sent both records to review. The officer releases them with one decision each. One Enron email of jokes that mentions California was called responsive.

Integrity. No run of the final pipeline left a withheld span recoverable. The check is not a formality: in test run 2 it failed a release in which the model had redacted words ("Lakeview Lodge") that also appear unredacted in other released records, which is exactly the leak it exists to catch.

History: what was changed after seeing test results

The test sets were meant to be run once. They were run four times, and three design faults were fixed in between; the table shows every run so the effect of seeing the test data is visible. No threshold was tuned on test data.

Run Pipeline change before it New held-out set written before it Responsiveness R PII recall Planted Precision Integrity
1 prompts final on the dev sets broadband, water, Enron 91.4% 81.0% (34/42) 84.0% 82.7% ok
2 request "including X, Y" treated as examples, not limits; consistency no longer overrides an official release park-concessions 86.0% 78.8% (41/52) 82.3% 88.7% failed (1 set)
3 "exempt or privileged records are still responsive"; parsed record types never limit; words the request uses are not withheld (non-privacy); consistency made privacy-protective again (a detail boxed anywhere is boxed everywhere) bus-contract 98.0% 91.7% (55/60) 93.0% 86.6% ok
4 medical spans must include the whole revealing phrase (motivated by the dev misses CP-006 and AP-05; the same pattern showed in the test runs) – 100% 96.7% (58/60) 97.2% 78.9% ok
5 a contact the model keeps as official is redacted unless on the agency's domains (or a staff/office number); a redacted person's first name standing alone is redacted in the same record – 100.0% 98.3% (59/60) 98.6% 83.1% ok

Run 5 (30 Sep, direct route to the same weights, foia-desk/test-run-5-scores.json) is the current pipeline: responsiveness precision / recall 96.2% / 100.0%, keep violations 3, 0 of 95 withheld spans recoverable, cost $0.77 per 1,000 records at list price (217 calls, 77 records). The page quotes this file.

  • Run 1 found that the model read "including reviewer evaluations, correspondence ..." as a limit and dropped a press release and a customer's letter.
  • Run 2 found that the model called emails with agency counsel "not responsive" because they were privileged, and that it withheld words the request itself uses, which then appeared unredacted elsewhere (the integrity failure).
  • The cleanest numbers are the sets that were new when a pipeline first met them: park-concessions in run 2 (responsiveness 10/10, PII 9/10, integrity failed on the request-words fault above) and bus-contract in run 3 (responsiveness 7/8, PII 5/8: one missed roster record took a licence number and a date of birth with it, and one medical phrase was partly covered).

Checks outside the model

  • The release PDF writer never writes withheld text; tests/test_foia.py checks that a box drawn over live text (the classic failure) is caught by the pdfplumber / filing-preflight check, and that digits split by layout are still found.
  • The signed record verifies at /record/verify and fails, naming the entry, when one exemption is changed (self-host run below). It holds no withheld text: the tests search it for the planted SSN, name and medical words.
  • Self-host (fresh clone into a clean directory, docker build, the assemble prompt's api service with a named volume, direct route to the already-running vLLM): coastal-permits end to end in 17 s, 41 calls, every disposition as expected, release check ok on 28 spans, the record verifies and the tampered copy fails at entry 4, one officer decision recorded and the ledger verifies.
  • Nightly smoke (scripts/smoke/foia-desk.py, three records): pass, 24 s, 6 signed receipts, $0.0025.

Limits of this eval

  • 48 synthetic held-out records and 29 public emails are a small sample; the labels are one AI reviewer's.
  • Synthetic records are cleaner than real ones: no OCR noise, no attachments, no long threads.
  • The hosted route (stated confidence) was measured; the self-host route (log-probabilities) was only smoke-tested.
  • Latency was measured on a shared, often busy gateway.