Skip to content
decosa

The full write-up behind the numbers on the tool’s page, as the team that built it wrote it: data, method, results and caveats. Internal names are removed; nothing else is edited.

Eval: report integrity (vertical 32)

Measured 25 Sep 2026 on our server, gateway route (Qwen3.8-27B NVFP4, every call receipted). Script: scripts/report_eval.py; data: decosa_api/verticals/report/data/cases.json (built by scripts/report_cases.py); raw rows and metrics in docs/evals/report-integrity/.

Data

12 invented incidents (5 police, 3 security, 2 EMS, 2 workplace), each with a timestamped transcript (13-22 lines), a faithful report (7-11 sentences, everything said on the recording, every key event covered) and a tampered report with six planted discrepancies:

  • 2 contradictions: a sentence rewritten so the recording says otherwise (a number, a refused consent becoming given, a denial becoming an admission, who said what);
  • 2 additions: a sentence the recording does not contain (an odor, bloodshot eyes, "he punched me", a prior history);
  • 2 omissions: a sentence about a key event removed (refused search, rights and a lawyer request, handcuffs too tight, a declined ambulance, risks of refusal explained).

The split is fixed: dev (traffic stop, retail detention, EMS fall, forklift) was used while writing the prompts; test (8 cases) was run once after the prompts were frozen. One prompt change came from dev: the coverage judge over-used "partly" for minor details (5/34 key events on faithful dev reports), so its definition was narrowed.

Caveats: the building agent wrote the scenarios and the plants, so they are clear-cut; real reports are messier, and a real test needs real reports against real footage, reviewed by a lawyer. One run of the test split.

Scoring (no model in the scoring)

  • addition or contradiction found: the planted sentence is flagged (partial, unsupported or contradicted); "exact" when a contradiction is marked contradicted;
  • omission found: a key event judged missing (or partly) sits on the planted transcript lines;
  • false flags: flagged sentences on faithful reports; false omissions: key events judged missing or partly on faithful reports.

Results

dev (4 cases, run 2) test (8 cases, held out) demo pair on ASR transcripts
additions found 8/8 16/16 4/4
contradictions found (exact) 8/8 (8/8) 16/16 (16/16) 4/4 (4/4)
omissions found, missing or partly (missing only) 8/8 (8/8) 16/16 (13/16) 4/4 (4/4)
faithful sentences flagged (hard: unsupported or contradicted) 2/39 (0) 0/64 (0) 1/19 (0)
unplanted sentences flagged in tampered reports 0/23 0/32 2/11
key events flagged on faithful reports, missing or partly (missing only) 2/40 (0) 9/68 (2/68) 2/20 (1/20)
verdicts that could not be made 0 0 0

Test split: 284 model calls, 284 receipts, 24,796 generated and 320,023 prompt tokens for 16 checks (about 18 calls, 1,550 generated and 20,000 prompt tokens per check with key events). Mean 83 s per check (43-168 s) while the gateway was shared with other evaluation jobs; dev checks earlier in the day took 8-36 s.

What the false flags were: "She signed the citation" marked partial (the recording has "where do I sign?" but not the signing: a fair catch); "wait in the office" partial (the recording says "come to the office" and "wait while I call"); on faithful reports, key events marked partly for a dropped instruction ("stop pulling") or a detail, and once a real tension in our own faithful report (the patient said both "just for a minute" and "maybe five minutes").

ASR run: the traffic-stop and bar-fight reports checked against the MOSS-Transcribe-Diarize transcripts of their synthetic audio, errors included. Plants found 12/12.

  • The audio changed on 26 Sep 2026. The first build used macOS system voices, which Apple licenses for personal use. Both recordings are now spoken by Decosa house voices (Kokoro-82M stock voicepacks, rendered by scripts/build_audio.py at the first build's turn times). Before a recording is written, the consent ledger's gate is asked for each voice (project decosa-report-demo, purpose character_dialogue). The entries are ce_79b931b17748 (af_heart), ce_524e0dbeda50 (am_michael), ce_80e655c7a3fb (am_fenrir), ce_6e39444e2d60 (af_nicole), ce_474026dce483 (bm_fable) and ce_9216ccd9b3b9 (am_puck); the decisions are in demo_scripts/voices.json.
  • Re-run on the new audio (re-diarized, run 1 again): plants still 12/12 (additions 4/4, contradictions 4/4 exact, omissions 4/4). Faithful sentences flagged: 2/19 (1/19 on the first build); flagged as contradicted: 1/19 (0/19); events judged missing on the faithful reports 3/20 (2/20).
  • Which errors. The first build's transcripts had "Silver on the Civic", "nine hundred eleven", "Police DMS" and "Cameron" for "Camera on"; the new ones have none of those. They have "Myra" for the bartender Mara once, and that is the new false flag: the faithful sentence "I asked the bartender, Mara, to call 911" was marked contradicted because the transcript has Vance telling "Myra" to call. The limit it shows is the same as before: the check trusts the transcript, so a mishearing decides the verdict. The writer asks the author to confirm likely mishearings.

Not measured

  • Real incident reports against real body-worn-camera audio (noise, overlap, radio traffic).
  • The first-draft writer's quality (the demo drafts were read by the building agent only).
  • The lite and best tiers.

Verdict

Would a buyer pay? A defence office or oversight board would pay for the check if it holds on real footage: it finds planted changes reliably, says where on the recording to look, rarely flags a faithful sentence, and signs the result. The agency side (draft, retained first draft, edit record, CA/UT disclosure) works end to end but competes with drafting tools bundled with the camera; it sells as "check before you sign" and as the audit trail California now requires. Missing before a sale: a pilot on real, consented footage with a lawyer scoring the output; video (what a camera sees is out of scope); a published diarizer image for self-host; and a speech-recognition quality number on body-cam audio.