Skip to content
decosa

18 · Legal · live

Deposition and hearing digest

Open the toolJSON

Eval results

Scored on a held-out or test splitRun 24 Sep 2026Eval write-up (decosa-api, access required)

  • Contradiction finder: planted conflicts found (held-out set)9 / 9 in both runstest splitn = 9
  • Contradiction finder: false-positive flags, run 1 / run 22 / 1 (precision 0.82 / 0.90)test splitn = 5All false flags were planted traps; 15 flags over 2 runs, very small
  • Digest sentences fully supported by their cited lines (hand-checked)33 / 40dev (tuned on)n = 407 partly supported (mostly a dropped hedge), 0 not supported; one annotator, not a lawyer
  • Cite checker: wrong cite flagged59 (100%)syntheticn = 59
  • Cite checker: changed fact flagged49 (91%)syntheticn = 54Unchanged sentences flagged on a second check: 3 of 60 (5%)

Dataset

Two public-domain congressional hearing excerpts (govinfo), a fictional two-witness demo pair used while writing the prompts, and a held-out fictional contradiction set (3 matters, 9 planted conflicts, 5 traps, 1 ambiguous pair).

Caveats

  • The held-out set was written by the same agent that wrote the prompts, before the eval ran.
  • Nothing was checked by a lawyer, and no real deposition was used.
  • Coverage (does the digest include everything important) is not measured; a 300-page deposition is not measured.
  • The audio path is tested only with a fake diarizer; its accuracy on deposition audio is unknown.
  • A number guard added after seeing the misses (51/54) is measured on the set it was designed from, so it is not held out.

Nightly smoke check

Loading the nightly status…

Result
pass
Run
30 Sep 2026
Latency, this run
n/a
p50 over passed runs
40 s
Receipts
48
Model calls
n/a
Tokens
n/a
Cost per run
$0.014

Self-host verification

Verified on 25 Sep 2026: fresh clone, api image built, the prompt's .env and compose used as written, sample against local model servers

The step 5 smoke passed as written in 23 s: every lane, 54 attested receipts, the 2:15 pm / 11:30 am conflict marked INCONSISTENT, the record verified and the Word export opened. A PDF transcript parsed as numbered (4:1-6:14), and with the diarize service /deposition/rough returned an uncertified rough transcript. Verified on 2026-09-25: images build, services start, sample passes end to end against local model servers equivalent to the documented ones; model-server startup itself not re-verified.

Rehearsal bundle: deposition.zip (4 KB, 9 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.

Known limits

  • Hosted numbers are the production smoke check of the sample, run 5 times in a row on 30 Sep 2026 (all passed); with 5 runs the slowest-1-in-20 figure is simply the slowest run.
  • The claim check shows each sentence is supported by the lines it cites; it does not show the digest covers everything important.
  • Hosted is for public-record and fictional transcripts only.

Models and licences

Standard tier. Licence posture: permissive (Apache, MIT or BSD).

All quality evidence

Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.

Lite · one 48 GB card (1)
  • cite-check accuracy on transcripts: not measured yet
Standard · the hosted demo, one 96 GB card (4)
  • hand check: sentences fully supported by their cited lines: 33/40decosa-api docs/evals/deposition.md, measured on our server 2026-09-24, gateway route; checked by the building agent, not a lawyer
  • cite checker: wrong cites / changed facts flagged: 59/59 / 49 of 54decosa-api docs/evals/deposition.md, measured on our server 2026-09-24, gateway route
  • contradictions, held-out fictional set: recall / precision: 9/9 / 0.82-0.90 (two runs)decosa-api docs/evals/deposition.md, measured on our server 2026-09-24, gateway route
  • cite check on real deposition transcripts (by a lawyer): not measured yet
Best · DeepSeek-V4-Flash on two more cards (1)
  • digest and check quality: not measured yet
Wanted · GLM-5.3-Flash on your own hardware (1)
  • digest quality: not measured yet

How we measure · All tools