18 · Legal · live
Deposition and hearing digest
Eval results
Scored on a held-out or test splitRun 24 Sep 2026Eval write-up (decosa-api, access required)
- Contradiction finder: planted conflicts found (held-out set)9 / 9 in both runstest splitn = 9
- Contradiction finder: false-positive flags, run 1 / run 22 / 1 (precision 0.82 / 0.90)test splitn = 5All false flags were planted traps; 15 flags over 2 runs, very small
- Digest sentences fully supported by their cited lines (hand-checked)33 / 40dev (tuned on)n = 407 partly supported (mostly a dropped hedge), 0 not supported; one annotator, not a lawyer
- Cite checker: wrong cite flagged59 (100%)syntheticn = 59
- Cite checker: changed fact flagged49 (91%)syntheticn = 54Unchanged sentences flagged on a second check: 3 of 60 (5%)
Dataset
Two public-domain congressional hearing excerpts (govinfo), a fictional two-witness demo pair used while writing the prompts, and a held-out fictional contradiction set (3 matters, 9 planted conflicts, 5 traps, 1 ambiguous pair).
Caveats
- The held-out set was written by the same agent that wrote the prompts, before the eval ran.
- Nothing was checked by a lawyer, and no real deposition was used.
- Coverage (does the digest include everything important) is not measured; a 300-page deposition is not measured.
- The audio path is tested only with a fake diarizer; its accuracy on deposition audio is unknown.
- A number guard added after seeing the misses (51/54) is measured on the set it was designed from, so it is not held out.
Nightly smoke check
Loading the nightly status…
- Result
- pass
- Run
- 30 Sep 2026
- Latency, this run
- n/a
- p50 over passed runs
- 40 s
- Receipts
- 48
- Model calls
- n/a
- Tokens
- n/a
- Cost per run
- $0.014
Self-host verification
Verified on 25 Sep 2026: fresh clone, api image built, the prompt's .env and compose used as written, sample against local model servers
The step 5 smoke passed as written in 23 s: every lane, 54 attested receipts, the 2:15 pm / 11:30 am conflict marked INCONSISTENT, the record verified and the Word export opened. A PDF transcript parsed as numbered (4:1-6:14), and with the diarize service /deposition/rough returned an uncertified rough transcript. Verified on 2026-09-25: images build, services start, sample passes end to end against local model servers equivalent to the documented ones; model-server startup itself not re-verified.
Rehearsal bundle: deposition.zip (4 KB, 9 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.
Known limits
- Hosted numbers are the production smoke check of the sample, run 5 times in a row on 30 Sep 2026 (all passed); with 5 runs the slowest-1-in-20 figure is simply the slowest run.
- The claim check shows each sentence is supported by the lines it cites; it does not show the digest covers everything important.
- Hosted is for public-record and fictional transcripts only.
Models and licences
Standard tier. Licence posture: permissive (Apache, MIT or BSD).
- Digest writer, cite checker and contradiction judgeQwen3.8-27B (NVIDIA NVFP4)Apache-2.0
- Self-host only: a recording to an uncertified rough transcript (POST /deposition/rough)MOSS-Transcribe-Diarize 0.9BApache-2.0
All quality evidence
Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.
Lite · one 48 GB card (1)
- cite-check accuracy on transcripts: not measured yet
Standard · the hosted demo, one 96 GB card (4)
- hand check: sentences fully supported by their cited lines: 33/40decosa-api docs/evals/deposition.md, measured on our server 2026-09-24, gateway route; checked by the building agent, not a lawyer
- cite checker: wrong cites / changed facts flagged: 59/59 / 49 of 54decosa-api docs/evals/deposition.md, measured on our server 2026-09-24, gateway route
- contradictions, held-out fictional set: recall / precision: 9/9 / 0.82-0.90 (two runs)decosa-api docs/evals/deposition.md, measured on our server 2026-09-24, gateway route
- cite check on real deposition transcripts (by a lawyer): not measured yet
Best · DeepSeek-V4-Flash on two more cards (1)
- digest and check quality: not measured yet
Wanted · GLM-5.3-Flash on your own hardware (1)
- digest quality: not measured yet