32 · Public sector · Legal · live
Report integrity
Eval results
Scored on a held-out or test splitRun 25 Sep 2026Eval write-up (decosa-api, access required)
- Planted additions found16/16test splitn = 16
- Planted contradictions found (exact)16/16 (16/16)test splitn = 16
- Planted omissions found, missing or partly (missing only)16/16 (13/16)test splitn = 16
- Faithful sentences flagged (hard: unsupported or contradicted)0/64 (0)test splitn = 64
- Key events flagged on faithful reports, missing or partly (missing only)9/68 (2/68)test splitn = 68
- Plants found on ASR transcripts of the demo pair12/12syntheticn = 12Two reports checked against MOSS-Transcribe-Diarize transcripts of synthetic audio, re-voiced 26 Sep 2026 with Decosa house voices (Kokoro-82M) and re-run: 2/19 faithful sentences flagged, one from a misheard name (first build: 12/12, 1/19).
Dataset
12 invented incidents (5 police, 3 security, 2 EMS, 2 workplace), each with a timestamped transcript, a faithful report and a tampered report with six planted discrepancies (2 contradictions, 2 additions, 2 omissions). Dev 4 cases, test 8 cases run once after the prompts were frozen.
Caveats
- The building agent wrote the scenarios and the plants, so they are clear-cut; real reports are messier.
- Synthetic incidents only: not measured on real incident reports against real body-worn-camera audio.
- One run of the test split.
- The check trusts the transcript: a mishearing the writer turns into a fact is marked as in the recording.
- The first-draft writer's quality and the lite and best tiers are not measured.
Nightly smoke check
Loading the nightly status…
- Result
- pass
- Run
- 25 Sep 2026
- Latency, this run
- n/a
- p50 over passed runs
- 97 s
- Receipts
- 21
- Model calls
- n/a
- Tokens
- n/a
- Cost per run
- $0.004
Self-host verification
Verified on 25 Sep 2026: fresh clone of decosa-api into a clean directory on our server, api image built from docker/api/Dockerfile, compose api service with a named data volume, direct route, local signing; torn down after
Verified on 2026-09-25: the image builds, the service starts, and both smoke tests in the assembly prompt pass end to end against a local model server equivalent to the documented one (the already-running Qwen3.8-27B vLLM on 127.0.0.1:8114, reached with network_mode host instead of the compose llm service); model-server startup itself not re-verified, and the diarizer path was not run. Check: 2 contradicted, 2 unsupported, refused search and the one-beer answer missing, 20 attested receipts, signature verified, 6 s. Draft: 12 sentences with two likely mishearings listed for the author, finalize 12 AI / 1 human, record verified, 13 s.
Rehearsal bundle: report-integrity.zip (4 KB, 13 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.
Known limits
- Hosted verification ran on the pre-release server (decosa-api the pre-release branch on our server, gateway route) driven from the branch site in headless Chromium, including 390 px; the production API gets this vertical when the branch merges.
- The check trusts the transcript. Speech-recognition errors pass through: in the demo the diarizer heard 'Camera on' as 'Cameron', and the draft named a bartender Cameron; the check marks it 'in the recording'. The writer now lists likely mishearings for the author to confirm.
- 'Not in the recording' is not 'false': what a camera cannot hear (smells, what someone saw) is flagged and needs the author's own account.
- Measured on 12 synthetic incidents written by the building agent, with clear-cut plants; real reports are messier. Not yet measured on real body-worn-camera audio or with a lawyer reviewing.
- Key-event coverage is noisier than the sentence check: 9 of 68 events on faithful reports were marked partly or missing, mostly detail a reader would not miss.
- Hosted runs took 80-130 s while the shared gateway was busy (8-36 s when quiet).
- Audio intake (/report/transcribe) is self-host only.
Models and licences
Standard tier. Licence posture: permissive (Apache, MIT or BSD).
- Key events, first-draft writer, sentence judge (the grounding checker) and event coverageQwen3.8-27B (NVIDIA NVFP4)Apache-2.0
- Recording to a timed, speaker-labelled transcript (POST /report/transcribe, self-host)MOSS-Transcribe-Diarize 0.9BApache-2.0
All quality evidence
Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.
Lite · one 48 GB card (1)
- planted discrepancies found / faithful sentences flagged: not measured yet
Standard · the hosted demo, one 96 GB card (4)
- held-out test, 8 synthetic incidents: planted additions / contradictions / omissions found: 16/16 / 16/16 / 16/16decosa-api docs/evals/report-integrity.md, measured on our server 2026-09-25, gateway route, one run; data written by the building agent, prompts frozen on a separate 4-case dev split
- faithful reports: sentences flagged / key events flagged as left out: 0/64 / 9/68 (2/68 as missing, 7 as partly)decosa-api docs/evals/report-integrity.md, measured on our server 2026-09-25, gateway route
- the two demo incidents on real speech-recognition transcripts: 4/4 additions, 4/4 contradictions, 4/4 omissions found; faithful sentences flagged 2/19 (1 partial, 1 contradicted)decosa-api docs/evals/report-integrity.md (asr run), measured on our server 2026-09-26, gateway route: the traffic-stop and bar-fight reports checked against MOSS-Transcribe-Diarize transcripts of their synthetic audio, re-voiced 26 Sep 2026 from macOS voices to Decosa house voices (Kokoro-82M) (the first build: 1/19 flagged, partial)
- real reports against real body-worn-camera audio, reviewed by a lawyer: not measured yet
Best · DeepSeek-V4-Flash on two more cards (1)
- planted discrepancies found / faithful sentences flagged: not measured yet
Wanted · two large judges from different families (1)
- planted discrepancies found / faithful sentences flagged: not measured yet