Eval: public-records desk (use case 34)
Run on 25 Sep 2026 on our server, hosted gateway route (Qwen3.8-27B NVFP4 through the model gateway, one call per typed
question with the stated confidence), on a GPU and gateway shared with other workloads' work. Scripts:
scripts/foia_samples.py (development sets), scripts/foia_testsets.py (held-out sets), scripts/foia_eval.py
(runs and scores). Score files: docs/evals/foia-desk/*-scores.json; the raw event logs are on our server in
~/data/foia-eval/runs/.
Data
| Set | Kind | Records | Planted spans | Notes |
|---|---|---|---|---|
| coastal-permits | dev (console sample), federal | 12 | 22 | resident complaints, SSNs, a deliberative email, agency counsel, a copy, out-of-range and off-topic records |
| alder-point | dev (console sample), CA CPRA | 9 | 14 | public comments with medical details, a draft memo, the city attorney |
| ferc-enron | dev (console sample), public | 10 | – | FERC-released Enron email, responsiveness only |
| broadband-grants | held out, federal | 14 | 26 | grants office; bank details, a sign-in sheet, an SSN form, a revised near-duplicate memo |
| water-rates | held out, CA CPRA | 14 | 24 | water district; shutoff list, medical details, a copy of a draft memo |
| park-concessions | held out, federal | 10 | 12 | written after test run 1 (see History) |
| bus-contract | held out, CA CPRA | 10 | 9 | written after test run 2 (see History) |
| ferc-enron-test | held out, public | 29 | – | responsiveness only, 11 responsive |
All synthetic names, numbers and domains are fictional (.example domains, 555-01xx phones). Labels (responsive or not, each span that should be redacted with its category, records to withhold in full, and "keep" text that must be released: official contacts, facts in deliberative records) were written by Claude (an AI agent) with the records, not by a records officer. The Enron emails were labelled by hand for a hypothetical request ("price caps or price mitigation in the California or western electricity markets, May to July 2001"); 9 ambiguous emails were dropped before labelling.
Metrics
- Responsiveness: a record counts as passed on when its disposition is anything but "not responsive"; precision and recall against the labels.
- Planted recall: a labelled span is redacted when redactions cover at least 80% of its characters, at every place it occurs in the record. "PII recall" is the same without the deliberative spans. Records labelled "withhold in full" are scored on the disposition and exemption instead.
- Exemption labels: for each redacted span, whether the exemption on the covering redaction is the one the label's category maps to in the set's list; for withheld records, whether the cited exemption matches.
- Redaction precision: the share of redactions in released records that overlap a labelled span. "Keep violations" are redactions over text labelled must-release.
- Integrity: the release PDF is read back by the built-in reader (every drawn string and every literal in the file), pdfplumber's text layer and the filing-preflight black-box check; a withheld span is "recovered" when its words (or, for numbers, its digits) appear in any reader's text.
- Pattern baseline: recall of the pattern finders alone, no model.
Final results (pipeline as shipped; test run 4)
| Held-out sets (77 records) | Dev sets (31 records) | |
|---|---|---|
| Responsiveness precision / recall | 98.1% / 100% (51 TP, 1 FP, 0 FN, 25 TN) | 95.5% / 100% |
| Planted spans redacted | 69 of 71 (97.2%) | 36 of 36 |
| Personal-data spans redacted (PII recall) | 58 of 60 (96.7%) | 30 of 30 |
| Pattern finders alone | 33 of 71 (46.5%) | 17 of 36 |
| Deliberative spans redacted | 11 of 11 | 6 of 6 |
| Exemption label correct on redacted spans | 69 of 69 | 36 of 36 |
| Records with counsel withheld in full, right exemption | 6 of 6 | 4 of 4 |
| Redaction precision / keep violations | 78.9% (71 of 90) / 3 | 94.7% / 0 |
| Release check: withheld spans recoverable | 0 of 102 (all readers; 0 characters under boxes) | 0 of 46 |
| Model calls / cost at list price | 216 calls, $0.059: $0.76 per 1,000 records | $0.94 per 1,000 |
| Wall time (shared gateway) | 178 s for 77 records | 157 s for 31 |
By kind (held out): names 19/19, phones 9/9, addresses 7/7, emails 6/6, account numbers 5/5, dates of birth 4/4, driver's licences 2/2, SSN 1/1, medical 5/7.
What it misses. Both misses are medical details: (1) WR-12, "our meter reader fractured his wrist on the job ... off on workers' compensation for six weeks": the worker is unnamed, and the model released the record in full, although coworkers could identify him; (2) WR-01, "My mother lives with me and uses an oxygen concentrator around the clock": 61% covered, the clause saying who lives there was left. Medical details are the weakest kind in every run (the model tends to box the diagnosis, not the words around it), so the console flags nothing extra for them today: an officer's read of each released record is still required.
Over-redaction. Most unlabelled redactions are company officials' names and business emails (a co-op's general manager, a concessionaire's manager, a bus company's contracts manager): the model called the person private in one record, and the consistency step, which is privacy-protective by design, then boxed the same words everywhere and sent both records to review. The officer releases them with one decision each. One Enron email of jokes that mentions California was called responsive.
Integrity. No run of the final pipeline left a withheld span recoverable. The check is not a formality: in test run 2 it failed a release in which the model had redacted words ("Lakeview Lodge") that also appear unredacted in other released records, which is exactly the leak it exists to catch.
History: what was changed after seeing test results
The test sets were meant to be run once. They were run four times, and three design faults were fixed in between; the table shows every run so the effect of seeing the test data is visible. No threshold was tuned on test data.
| Run | Pipeline change before it | New held-out set written before it | Responsiveness R | PII recall | Planted | Precision | Integrity |
|---|---|---|---|---|---|---|---|
| 1 | prompts final on the dev sets | broadband, water, Enron | 91.4% | 81.0% (34/42) | 84.0% | 82.7% | ok |
| 2 | request "including X, Y" treated as examples, not limits; consistency no longer overrides an official release | park-concessions | 86.0% | 78.8% (41/52) | 82.3% | 88.7% | failed (1 set) |
| 3 | "exempt or privileged records are still responsive"; parsed record types never limit; words the request uses are not withheld (non-privacy); consistency made privacy-protective again (a detail boxed anywhere is boxed everywhere) | bus-contract | 98.0% | 91.7% (55/60) | 93.0% | 86.6% | ok |
| 4 | medical spans must include the whole revealing phrase (motivated by the dev misses CP-006 and AP-05; the same pattern showed in the test runs) | – | 100% | 96.7% (58/60) | 97.2% | 78.9% | ok |
| 5 | a contact the model keeps as official is redacted unless on the agency's domains (or a staff/office number); a redacted person's first name standing alone is redacted in the same record | – | 100.0% | 98.3% (59/60) | 98.6% | 83.1% | ok |
Run 5 (30 Sep, direct route to the same weights, foia-desk/test-run-5-scores.json) is the current pipeline: responsiveness precision / recall 96.2% / 100.0%, keep violations 3, 0 of 95 withheld spans recoverable, cost $0.77 per 1,000 records at list price (217 calls, 77 records). The page quotes this file.
- Run 1 found that the model read "including reviewer evaluations, correspondence ..." as a limit and dropped a press release and a customer's letter.
- Run 2 found that the model called emails with agency counsel "not responsive" because they were privileged, and that it withheld words the request itself uses, which then appeared unredacted elsewhere (the integrity failure).
- The cleanest numbers are the sets that were new when a pipeline first met them: park-concessions in run 2 (responsiveness 10/10, PII 9/10, integrity failed on the request-words fault above) and bus-contract in run 3 (responsiveness 7/8, PII 5/8: one missed roster record took a licence number and a date of birth with it, and one medical phrase was partly covered).
Checks outside the model
- The release PDF writer never writes withheld text;
tests/test_foia.pychecks that a box drawn over live text (the classic failure) is caught by the pdfplumber / filing-preflight check, and that digits split by layout are still found. - The signed record verifies at
/record/verifyand fails, naming the entry, when one exemption is changed (self-host run below). It holds no withheld text: the tests search it for the planted SSN, name and medical words. - Self-host (fresh clone into
a clean directory,docker build, the assemble prompt's api service with a named volume, direct route to the already-running vLLM): coastal-permits end to end in 17 s, 41 calls, every disposition as expected, release check ok on 28 spans, the record verifies and the tampered copy fails at entry 4, one officer decision recorded and the ledger verifies. - Nightly smoke (
scripts/smoke/foia-desk.py, three records): pass, 24 s, 6 signed receipts, $0.0025.
Limits of this eval
- 48 synthetic held-out records and 29 public emails are a small sample; the labels are one AI reviewer's.
- Synthetic records are cleaner than real ones: no OCR noise, no attachments, no long threads.
- The hosted route (stated confidence) was measured; the self-host route (log-probabilities) was only smoke-tested.
- Latency was measured on a shared, often busy gateway.