173 · Legal · live
Settlement video from the case file
Eval results
Scored on a held-out or test splitRun 29 Sep 2026Eval write-up (decosa-api, access required)
- Cites on the right page, rendered lines376 / 376test splitn = 3766 held-out synthetic matters; prompts and checks frozen before
- Dates in rendered lines right53 / 53test splitn = 53
- Cites on the right page, matters written blind by another agent143 / 144held outn = 144139 / 144 by the writer's answer key; on review, four of the five unmatched cites were on the right page (findings the key did not list)
- Cites on the right page, post-fix held-out matters201 / 201test splitn = 2013 new seeds run once after two fixes made on the first held-out run
- Every itemized charge read and summed to the cent11 / 11 matterstest splitn = 11the default total also leaves out charges with no matching visit; it equals the key in 7 / 11 because the chronology missed 1-3 visits in four matters
- Planted bill problems flagged43 / 44test splitn = 44before the injury 11/11, duplicate 11/11, no matching visit 11/11 (7 false flags), statement total off 10/11 (one total on a scan was unreadable and said so)
- Rendered lines a blind judge found supported by their cites77 / 80test splitn = 803 partly, 0 not supported, 0 misleading; Claude Code Opus 5.5, blind, saw at most six cites per line
- Lines held back for the lawyer13 / 135test splitn = 1354 right to hold, 3 dates outside what the line cites, 6 over-strict (3 from a pronoun issue since fixed; 0 over-strict in the post-fix split)
Dataset
6 held-out and 3 post-fix synthetic matters from the chronology generator (records, itemized bills with planted problems, a signed client statement), plus 2 matters written blind by another agent (38 pages); each run end to end: chronology, the client's recorded consent, then the draft.
Caveats
- Synthetic matters only; the generated ones share an author and a template family with the tool.
- The blind matters were written by another agent but rendered with the same page renderer.
- Answer keys are page-level; box placement comes from the chronology, whose own eval measured it.
- The lawyer's review time is an estimate, not timed with a lawyer.
- Two fixes were made after the first held-out run (the judge's statement wording, re-reads of unsure bill regions); the post-fix split and the blind split were run once after them.
Nightly smoke check
Loading the nightly status…
- Result
- pass
- Run
- 30 Sep 2026
- Latency, this run
- n/a
- p50 over passed runs
- 104 s
- Receipts
- 23
- Model calls
- n/a
- Tokens
- n/a
- Cost per run
- $0.010
Self-host verification
Verified on 29 Sep 2026: fresh clone of the branch into a clean directory, run with the host's Python environment (no container build: the server's root disk was full at the time), direct route to the local Qwen3.8-27B, the running document reader, local signing, real-matters mode on
The rehearsal bundle passed 11/11 in 38.5 s (draft, tie-out flags, render, record verified and failed when changed) and the smoke module passed in 19.9 s.
Rehearsal bundle: settlement-video.zip (2 KB, 11 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.
Known limits
- Hosted numbers are the whole sample task measured on production (draft + one line re-checked + render with a house voice, through the production API), run 5 times in a row on 30 Sep 2026 (all passed); with 5 runs the slowest-1-in-20 figure is simply the slowest run.
- Measured on synthetic matters only; real records, bills and photos are not measured, and the lawyer's review time (10-20 minutes estimated) has not been timed with a lawyer.
- When the chronology misses a visit, a real charge on that day is flagged as having no matching visit and left out of the default total until the lawyer puts it back (4 of 11 eval matters, $186-594).
- A printed statement total on a degraded scan or fax is sometimes unreadable; the tool says so and the video uses the sum of the lines.
- A child's photo is refused until a guardian consent path for this purpose is approved.
- No causation, future-care, wage-loss or billed-versus-paid figures: only what the chronology and the itemized bills hold.
Models and licences
Standard tier. Licence posture: permissive (Apache, MIT or BSD).
- The script checks (refs, cites attached in code, numbers, spinal levels and doses), the bills tie-out, the scene plan, the frames and the cite sheet (no model; CPU)decosa-api settlement video (decosa_api/verticals/settlement), importing the grounding, numeric-grounding, consent-ledger, provenance and record blocksAGPL-3.0-or-later
- The 7-scene script (one call: which chronology entries and statement paragraphs each line rests on) and one grounding verdict per line against the cited record text; re-reads of bill and statement regions the parser was unsure ofQwen3.8-27B (NVFP4)Apache-2.0
- Finds the regions of each bill and statement page (tables, text) with their boxesDocling 2.130 with the Heron layout model (document reader block)MIT (Docling) + Apache-2.0 (weights)
- Reads the bill tables as cells and the statement's paragraphsPaddleOCR-VL-1.6 (0.9B, document reader block)Apache-2.0
- The stock house voice that reads the approved script, when the lawyer picks itKokoro-82M (stock voicepacks am_michael, af_heart)Apache-2.0
- The attorney's own voice, generated from the attorney's consent recording, only with an active consent-ledger entry for this matterChatterbox MultilingualMIT
All quality evidence
Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.
Lite · captions or your own recording, no voice models (2)
- Script, checks and tie-out: same as standard (the voice does not change them)decosa-api docs/evals/settlement-video.md
- Render time, captions only: not measured separately; frames and encoding took 41-43 s of the 60-68 s rendersdecosa-api docs/evals/settlement-video.md, render table
Standard · reader, model, house voices and the consented clone (hosted demo) (6)
- Cites on the right page, lines that render (6 held-out synthetic matters): 376 / 376decosa-api docs/evals/settlement-video.md, held-out seeds, 29 Sep 2026
- Cites on the right page, 2 matters written blind by another agent: 143 / 144 on review (139 / 144 by the writer's key)decosa-api docs/evals/settlement-video.md, blind split
- Dates in rendered lines that match the answer key (all splits): 100 / 100decosa-api docs/evals/settlement-video.md: 53 held out, 17 blind, 30 post-fix
- Every itemized charge read and summed to the cent: 11 / 11 mattersdecosa-api docs/evals/settlement-video.md: the default total also leaves out charges with no matching visit, which equals the key in 7 / 11 (the chronology missed 1-3 visits in the others)
- Planted bill problems flagged (before the injury, duplicate, no matching visit, statement total off): 11/11, 11/11, 11/11, 10/11decosa-api docs/evals/settlement-video.md, all splits
- Rendered lines a blind judge found supported by their cited text: 77 / 80 (3 partly, 0 not supported)decosa-api docs/evals/settlement-video.md, Claude Code Opus 5.5 blind, held-out sample