Skip to content
decosa

32 · Public sector · Legal · live

Report integrity

Open the toolJSON

Eval results

Scored on a held-out or test splitRun 25 Sep 2026Eval write-up (decosa-api, access required)

  • Planted additions found16/16test splitn = 16
  • Planted contradictions found (exact)16/16 (16/16)test splitn = 16
  • Planted omissions found, missing or partly (missing only)16/16 (13/16)test splitn = 16
  • Faithful sentences flagged (hard: unsupported or contradicted)0/64 (0)test splitn = 64
  • Key events flagged on faithful reports, missing or partly (missing only)9/68 (2/68)test splitn = 68
  • Plants found on ASR transcripts of the demo pair12/12syntheticn = 12Two reports checked against MOSS-Transcribe-Diarize transcripts of synthetic audio, re-voiced 26 Sep 2026 with Decosa house voices (Kokoro-82M) and re-run: 2/19 faithful sentences flagged, one from a misheard name (first build: 12/12, 1/19).

Dataset

12 invented incidents (5 police, 3 security, 2 EMS, 2 workplace), each with a timestamped transcript, a faithful report and a tampered report with six planted discrepancies (2 contradictions, 2 additions, 2 omissions). Dev 4 cases, test 8 cases run once after the prompts were frozen.

Caveats

  • The building agent wrote the scenarios and the plants, so they are clear-cut; real reports are messier.
  • Synthetic incidents only: not measured on real incident reports against real body-worn-camera audio.
  • One run of the test split.
  • The check trusts the transcript: a mishearing the writer turns into a fact is marked as in the recording.
  • The first-draft writer's quality and the lite and best tiers are not measured.

Nightly smoke check

Loading the nightly status…

Result
pass
Run
25 Sep 2026
Latency, this run
n/a
p50 over passed runs
97 s
Receipts
21
Model calls
n/a
Tokens
n/a
Cost per run
$0.004

Self-host verification

Verified on 25 Sep 2026: fresh clone of decosa-api into a clean directory on our server, api image built from docker/api/Dockerfile, compose api service with a named data volume, direct route, local signing; torn down after

Verified on 2026-09-25: the image builds, the service starts, and both smoke tests in the assembly prompt pass end to end against a local model server equivalent to the documented one (the already-running Qwen3.8-27B vLLM on 127.0.0.1:8114, reached with network_mode host instead of the compose llm service); model-server startup itself not re-verified, and the diarizer path was not run. Check: 2 contradicted, 2 unsupported, refused search and the one-beer answer missing, 20 attested receipts, signature verified, 6 s. Draft: 12 sentences with two likely mishearings listed for the author, finalize 12 AI / 1 human, record verified, 13 s.

Rehearsal bundle: report-integrity.zip (4 KB, 13 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.

Known limits

  • Hosted verification ran on the pre-release server (decosa-api the pre-release branch on our server, gateway route) driven from the branch site in headless Chromium, including 390 px; the production API gets this vertical when the branch merges.
  • The check trusts the transcript. Speech-recognition errors pass through: in the demo the diarizer heard 'Camera on' as 'Cameron', and the draft named a bartender Cameron; the check marks it 'in the recording'. The writer now lists likely mishearings for the author to confirm.
  • 'Not in the recording' is not 'false': what a camera cannot hear (smells, what someone saw) is flagged and needs the author's own account.
  • Measured on 12 synthetic incidents written by the building agent, with clear-cut plants; real reports are messier. Not yet measured on real body-worn-camera audio or with a lawyer reviewing.
  • Key-event coverage is noisier than the sentence check: 9 of 68 events on faithful reports were marked partly or missing, mostly detail a reader would not miss.
  • Hosted runs took 80-130 s while the shared gateway was busy (8-36 s when quiet).
  • Audio intake (/report/transcribe) is self-host only.

Models and licences

Standard tier. Licence posture: permissive (Apache, MIT or BSD).

All quality evidence

Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.

Lite · one 48 GB card (1)
  • planted discrepancies found / faithful sentences flagged: not measured yet
Standard · the hosted demo, one 96 GB card (4)
  • held-out test, 8 synthetic incidents: planted additions / contradictions / omissions found: 16/16 / 16/16 / 16/16decosa-api docs/evals/report-integrity.md, measured on our server 2026-09-25, gateway route, one run; data written by the building agent, prompts frozen on a separate 4-case dev split
  • faithful reports: sentences flagged / key events flagged as left out: 0/64 / 9/68 (2/68 as missing, 7 as partly)decosa-api docs/evals/report-integrity.md, measured on our server 2026-09-25, gateway route
  • the two demo incidents on real speech-recognition transcripts: 4/4 additions, 4/4 contradictions, 4/4 omissions found; faithful sentences flagged 2/19 (1 partial, 1 contradicted)decosa-api docs/evals/report-integrity.md (asr run), measured on our server 2026-09-26, gateway route: the traffic-stop and bar-fight reports checked against MOSS-Transcribe-Diarize transcripts of their synthetic audio, re-voiced 26 Sep 2026 from macOS voices to Decosa house voices (Kokoro-82M) (the first build: 1/19 flagged, partial)
  • real reports against real body-worn-camera audio, reviewed by a lawyer: not measured yet
Best · DeepSeek-V4-Flash on two more cards (1)
  • planted discrepancies found / faithful sentences flagged: not measured yet
Wanted · two large judges from different families (1)
  • planted discrepancies found / faithful sentences flagged: not measured yet

How we measure · All tools