Skip to content
decosa

The full write-up behind the numbers on the tool’s page, as the team that built it wrote it: data, method, results and caveats. Internal names are removed; nothing else is edited.

Eval: Make a settlement video from the case file (use case 173)

29 Sep 2026. Code: decosa_api/verticals/settlement/; runner: scripts/settlement_eval.py; sample builder: scripts/settlement_samples.py. All matters are synthetic; no real person, record, bill or photo is used.

What is measured, and how

A settlement video is only as good as its claims. The eval asks, for every line that would appear in the video:

  1. Cite on the right page. Every cite (exhibit, page, box, attached in code from the chronology entry, the statement paragraph or the bill row a line rests on) must point at a page where the answer key has that fact.
  2. Dates right. Every date a rendered line states must be the date of a fact it cites in the answer key (or the date of injury).
  3. Bills exact. The total the video shows must equal the answer key to the cent, and every planted bill problem must be flagged: a charge before the injury, a charge billed twice, a charge with no matching visit, and a statement whose printed total is not the sum of its own lines.
  4. Support, judged blind. A sample of rendered lines with their cited record text, judged by a model that saw only those inputs (Claude Code Opus 5.5, blind, n = 80).
  5. What is held back. Lines the checks mark (the lawyer fixes them or marks them as their own words), and why.

Each matter runs end to end on the pre-release server over the gateway route: POST /chronology/run on the record files, the client's consent enrolled through POST /settlement/consent (a consent sentence read by a stock voice for the fictional client, transcribed and scored), then POST /settlement/draft. Scoring is code against the answer key.

Splits.

  • Dev: the Whitlock sample (seed 7) and seed 999. Prompts and checks were written and fixed on these.
  • Held out (frozen): seeds 1101, 1202, 1303, 1404, 1505, 1606, run once after the prompts and checks were frozen.
  • Blind: two matters written by a separate agent from a format brief only (no code, no prompts): 38 pages of records, bills and statements in different providers, injuries, mechanisms and wording. Run once, after two fixes made on what the held-out run showed (below).
  • Post-fix held out: seeds 2101, 2202, 2303, run once after those fixes.

The two fixes after the held-out run, stated so the numbers can be read correctly: (1) the statement paragraph shown to the grounding judge no longer carries a pronoun of its own (it had made the judge hold back correct lines that said "he"); (2) bill and statement pages now get the document reader's model re-read of regions the parser was unsure of, as the chronology does (a scanned statement total had been read as "___").

Results

Split Matters Cites on the right page Dates right Bills total exact Lines drafted / rendered / held back
Held out 6 376 / 376 53 / 53 6 / 6 135 / 122 / 13
Blind-written 2 139 / 144 by the key; 143 / 144 on review 17 / 17 2 / 2 46 / 40 / 6
Post-fix held out 3 201 / 201 30 / 30 3 / 3 66 / 61 / 5

Blind review of the five cites the key did not match: four point at the right page (two findings in a wrist MRI report and an operative report that the blind key did not list as events, and an operation the chronology labelled as a consultation, cited to its own operative report); one is partial (the line says the truck "struck the back" of the car, the cited text says "collision with pickup truck").

Planted bill problems flagged

Problem Held out Blind Post-fix
Charge before the injury (left out of the total) 6 / 6 2 / 2 3 / 3
Duplicate charge (counted once) 6 / 6 2 / 2 3 / 3
Charge with no matching visit 6 / 6 2 / 2 3 / 3
Statement total off its own lines 5 / 6 2 / 2 3 / 3
"Total not read" notices on other statements 2 1 1

"Bills total exact" in the first table is the sum of every readable charge, less charges before the injury and the second copy of a duplicate: 11 / 11 matters to the cent, every itemized line read right. After the blind reviews (below), the default changed to also leave out charges with no matching visit in the chronology, with each one listed for the lawyer to put back in one click. Re-scored on the same eleven drafts, that default total equals the key in 7 / 11: the other four leave out one to three real charges ($186-594) whose visits the chronology missed. The no-matching-visit flag caught all 11 planted charges and raised 7 false flags, all from visits the chronology did not find (it counts a day as care when any entry is dated that day).

The held-out miss: the statement whose printed total was wrong was a scan, and the reader read its total as "___"; the tool said it could not check that statement against its total, and the video's figure (the sum of the lines) was still exact. "Total not read" notices are honest: on a degraded scan or fax the printed total could not be read, so that one comparison is left to the lawyer.

Blind support judgments (n = 80 rendered lines from the held-out split): 77 supported, 3 partly, 0 not supported, 0 misleading. The three partial lines: a therapy date range whose end date came from a cite beyond the six shown to the judge; "the emergency department diagnosed" when the cited text is the hospital's diagnosis list; "on that day" (the date of injury) where the cited text shows only the date.

Post-fix held out (3 matters): 5 of 66 lines held back, all right to hold (a mechanism line citing an entry that does not state it, two "L4-15" read slips, an invented "his sister", a therapy range ending outside what the line cites); no over-strict holds.

Lines held back in the held-out split (13 of 135). 4 were right to hold (a spinal-level read slip "L4-15" that the chronology carried from a scanned MRI page; a detail the statement does not say, "her friends"; a mechanism line citing an entry that does not state it; a wrong "last visit" date). 3 stated a date outside what the line cites (a therapy range ending on the discharge date without citing the discharge note). 6 were over-strict partial verdicts, 3 of them from the pronoun issue fixed above. None of the held-back lines would have reached the video without the lawyer's edit.

Time and cost (gateway, list price $0.30 / $1.50 per million tokens).

Step Measured
Draft (script and a check per line), held-out matters 16.6-31.0 s, 15-18 model calls, $0.0072-0.0083
Draft, blind matters 28.9-30.7 s, $0.0082-0.0098
Render with a house voice, the Whitlock sample (3:17 to 3:56 videos, 4 runs) 61.3-69.2 s: narration about 17-20 s (Kokoro on CPU), frames and encoding about 42-48 s on 32 cores
Render, captions only (browser end-to-end run, 2:32 video) 26.5 s
The chronology the video starts from (its own tool) 25.8-99.9 s per 16-page matter
The lawyer's review (estimate, not measured with a lawyer) 10-20 minutes for about 20-25 lines: read each line, open the cites of anything that looks off, fix or remove the 1-4 held-back lines, check the tie-out flags, approve 7 scenes

Blind reviews by two personas (Claude Code Opus 5.5 sub-agents; frames, script and cite sheet only)

  • Plaintiff's attorney, first version: would not show it as is; would with changes, on mid-value cases ($50-250k) where a commissioned film is not worth it; $150-300 a video once fixed. Complaints, most fixed before the second version: record pages showed adverse text and printed totals next to the video's figures (now the page recedes around the cited box and a provider's charges are cropped to their rows); gap bars on the timeline (now off by default); a no-record charge kept in the total (now left out by default); "Dana says" hedging (now third person, marked as her statement); thin "life now" (now every statement paragraph about it, including her closing wish); a cane line over a pill-organizer photo (photos now match the line or give way to a quote card); the bills counter showing a partial sum under the total's caption (fixed); "Made with Decosa" on the exhibit (removed). Not addressed: the client on camera or in her own recorded voice, causation and future-care language, billed versus paid.
  • Plaintiff's attorney, second version: would show it with changes, "now small: about an hour of my editing"; not as is. Confirmed fixed: the gap bar, the total matching the narration, the no-record charge out of the total, no "Dana says", the MRI box on the impression, reconciled counts, the end card. Still wrong then, and changed for the third version: the page outside the cited box was dimmed but readable (now covered, so only the cited words show), a bill page showed the excluded rows and the printed total (now only the included rows show, and "left out" moved to the cite sheet), "they" for the client (the matter now carries pronouns), the injuries scene opening with strain and sprain (now most serious first), three medication lines (now one). Not addressed: the client on camera, causation, future care, permanency and wages. Price once fixed: $200-350 a video, or $400-600 a month for a four-lawyer firm; as is, a trial.
  • Insurance adjuster, second version: "It doesn't move my number", but it raises confidence in the specials: the drops of the 2022 visit, the duplicate and the July charge with no visit note, and the use of line items over the misprinted total, are credited to counsel. Better than a typical demand package ("faster to verify, the math isn't inflated, and it flags its own exclusions"), weaker emotionally than a produced documentary. Pushback: the gap after surgery, the degenerative X-ray finding, no endpoint note (MMI, restrictions), billed not paid.

Caveats

  • Synthetic matters only. The generated split and the tool share an author and one template family; the blind split was written by another agent but rendered with the same page renderer. Real records, real bills and real client photos are not measured.
  • The answer keys are page-level. Box-level placement is checked by construction (cites come from the chronology, whose own eval measured boxes), not re-measured here.
  • The lawyer's review time is an estimate; no practising lawyer has timed it.
  • The blind support judge saw at most six cites per line.

Expected properties of the sample run (rehearsal bundle rehearsal/settlement-video/)

  1. The bills tie to $24,331.00 in code, with the $18 statement error, the duplicate, the 2022 pre-injury charge and the no-record charge flagged (the last three left out of the total).
  2. Every line that passes carries an exhibit, page and box cite.
  3. A line naming the scanned MRI's level "L4-15" is held back as not a valid spinal level.
  4. The render refuses until every scene is approved and every line passes or is marked as the attorney's own words.
  5. The MP4 carries a C2PA credential naming it a composite drawn in code, and the signed record verifies (and fails when changed).