Deposition and hearing digest: eval note (24 Sep 2026)
Small and honest. Everything below ran on our server through the model gateway (Qwen3.8-27B NVFP4, every call receipted),
on branch the pre-release branch. Scripts: scripts/deposition_client.py (saves a run), scripts/deposition_eval.py
(contradictions, mutations, handsheet). No GPU model was loaded for this work.
Data
- Public domain. Two U.S. congressional hearing excerpts from govinfo, U.S. Government works (17 U.S.C. 105):
S. Hrg. 118-216 (Senate Banking, 28 Mar 2023, CHRG-118shrg54561, printed pp. 21-24) and House Financial Services
Serial No. 118-12 (29 Mar 2023, CHRG-118hhrg52390, printed pp. 10-13). The PDFs have page numbers but no line numbers,
so page = printed page and line = the nth text line on that page from
pdftotext -layout.scripts/deposition_sources.pyrebuilds them and records each PDF's sha256 insamples/samples.json. - Fictional. A two-witness deposition pair written for the demo (warehouse fall; 4 planted conflicts). Used while writing the prompts, so it is not a test set.
- Held-out contradiction set.
tests/fixtures/deposition_eval/: three fictional matters, two witnesses each, in three layouts (numbered, plain without numbers, condensedpage:line), with 9 planted conflicts, 5 traps (look like conflicts, are not) and 1 ambiguous pair intruth.json. Written by the same agent that wrote the prompts, before the eval ran; the prompts were not changed after seeing its results.
1. Cite accuracy, hand-checked (40 sentences)
40 digest sentences sampled at random (seed 11) from one hearing run and one fictional run; each read against its cited lines by the building agent (Claude), not a lawyer, one annotator.
| count | |
|---|---|
| fully supported by the cited lines | 33 / 40 |
| partly supported (dropped hedge such as "I think" or "my understanding is", cite too narrow, an added "because") | 7 / 40 |
| not supported | 0 / 40 |
| cite checker agrees with the hand label (supported vs flagged) | 36 / 40 |
| of the 7 partial, flagged by the checker | 4 |
| of the 33 supported, flagged by the checker (false flags) | 1 |
The common digest error is a dropped hedge; the checker catches it about half the time.
2. Cite checker on programmatic errors
From 60 checked sentences (both runs) the script makes: a wrong cite (same length, 25+ lines away), and a changed fact (a number changed, a negation removed, or a speaker's name swapped). Each is judged the way the pipeline judges.
| case | n | flagged (PARTIAL or UNSUPPORTED) |
|---|---|---|
| wrong cite | 59 | 59 (100%) |
| changed fact | 54 | 49 (91%): numbers 26/28, speaker 17/18, negation 6/8 |
| unchanged sentence, second check | 60 | 3 flagged (5%) |
After seeing the misses, a fixed number guard was added (a number in the sentence that appears nowhere in the cited lines or their context turns SUPPORTED into PARTIAL). Applied offline to the same rows it catches 2 more (51/54) and flags 0 of the 61 supported sentences in the two runs. That is measured on the set it was designed from, so it is not a held-out number.
3. Contradiction finder, held-out set
| run | planted found | false-positive flags | precision | recall |
|---|---|---|---|---|
| 1 | 9 / 9 | 2 (both traps: "notice emailed" vs "I don't believe I received one"; "it's Exhibit 14" vs "never seen this email") | 0.82 | 1.00 |
| 2 | 9 / 9 | 1 (the notice trap) | 0.90 | 1.00 |
Run 1 was first scored with a matcher that assigned hits to the wrong planted item when facts sat within two lines of each other; the table uses the fixed matcher (closest planted phrase, both sides within 2 lines). 15 flags over 2 runs; very small. On the demo pair (not held out) 4/4 planted were found with 0 false flags in the 3 runs after a speaker-attribution fix (one of them flagged the leak conflict twice); the first run, before that fix, found 2. On the two real hearing excerpts nothing was flagged inconsistent (no known conflict there; not scored).
Known failure seen in the browser bug hunt (a 5-line hand-made pair, not scored): two unnamed witnesses where one "did not recall" a ladder and the other saw one were flagged inconsistent, and a pair whose context lines carried a neighbouring conflict was flagged. The judge now names each unnamed witness by transcript; the second case remains.
4. Speed and tokens (demo budget is 20,000 generated tokens)
| input | lines | sentences | pairs | wall time | generated tokens |
|---|---|---|---|---|---|
| fictional pair | 117 | 32-38 | 12 | 34-76 s | about 5,400 |
| hearing pair | 429 | 34-43 | 12 | 55-74 s | about 5,200-6,700 |
A 300-page deposition is not measured; dense Q/A ran at about 47 generated tokens per line, hearings at about 13.
What this does not show
- Nothing here was checked by a lawyer, and no real deposition was used (they are rarely public-domain).
- Coverage (does the digest include everything important) is not measured.
- The audio path (
/deposition/rough) is tested only with a fake diarizer; its accuracy on deposition audio is unknown.