{"schema_version":"1","site":"https://decosa.ai","id":"report-integrity","num":"32","name":"Report integrity","tool_name":"Check a police report against bodycam","short":"Report integrity","blurb":"Checks an incident report (police, security, EMS, workplace) against the transcript of its recording, sentence by sentence, and lists key events the report leaves out, with times you can cue up. Or drafts the report from the recording, keeps a signed first draft, and records which sentences a person changed before signing.","status":"live","labels":{"industry":["public-sector","legal"],"job":["review","attest"],"input":["text","voice"],"deploy":["selfhost"],"status":"live","output":["record","text"],"data":["pii","confidential"],"hardware":"gpu-96","licence":"permissive"},"industries":["public-sector","legal"],"runs_in":["selfhost"],"part_of":[],"built_from":["diarize","grounding","signed-record"],"models":"MOSS-Transcribe-Diarize · Qwen3.8-27B","where":"Self-host (hosted demo: synthetic material only)","hardware":"1× RTX PRO 6000 (96 GB) for Qwen3.8-27B and the diarizer; the check itself needs no GPU","final_artifact":"Every sentence marked against the recording with its times, the key events the report leaves out, and a signed statement; in draft mode a signed first draft and an edit record.","self_host_first":true,"verification":{"receipt_coverage":"full","summary":"Receipt per model call; signed check; signed first draft and hash-chained edit record","manual_qa":{"hosted":{"date":"2026-09-25","result":"pass","p50_ms":97000,"p95_ms":null,"runs":null,"receipts_per_run":21,"cost_per_run_usd":0.004},"selfhost":{"date":"2026-09-25","result":"pass","method":"fresh clone of decosa-api into a clean directory on our server, api image built from docker/api/Dockerfile, compose api service with a named data volume, direct route, local signing; torn down after","notes":"Verified on 2026-09-25: the image builds, the service starts, and both smoke tests in the assembly prompt pass end to end against a local model server equivalent to the documented one (the already-running Qwen3.8-27B vLLM on 127.0.0.1:8114, reached with network_mode host instead of the compose llm service); model-server startup itself not re-verified, and the diarizer path was not run. Check: 2 contradicted, 2 unsupported, refused search and the one-beer answer missing, 20 attested receipts, signature verified, 6 s. Draft: 12 sentences with two likely mishearings listed for the author, finalize 12 AI / 1 human, record verified, 13 s."},"known_limits":["Hosted verification ran on the pre-release server (decosa-api the pre-release branch on our server, gateway route) driven from the branch site in headless Chromium, including 390 px; the production API gets this vertical when the branch merges.","The check trusts the transcript. Speech-recognition errors pass through: in the demo the diarizer heard 'Camera on' as 'Cameron', and the draft named a bartender Cameron; the check marks it 'in the recording'. The writer now lists likely mishearings for the author to confirm.","'Not in the recording' is not 'false': what a camera cannot hear (smells, what someone saw) is flagged and needs the author's own account.","Measured on 12 synthetic incidents written by the building agent, with clear-cut plants; real reports are messier. Not yet measured on real body-worn-camera audio or with a lawyer reviewing.","Key-event coverage is noisier than the sentence check: 9 of 68 events on faithful reports were marked partly or missing, mostly detail a reader would not miss.","Hosted runs took 80-130 s while the shared gateway was busy (8-36 s when quiet).","Audio intake (/report/transcribe) is self-host only."],"nightly_covers":null},"nightly":"https://api.decosa.ai/verify/status"},"eval_summary":{"metrics":[{"name":"Planted additions found","value":"16/16","unit":null,"n":16,"split":"test","note":null},{"name":"Planted contradictions found (exact)","value":"16/16 (16/16)","unit":null,"n":16,"split":"test","note":null},{"name":"Planted omissions found, missing or partly (missing only)","value":"16/16 (13/16)","unit":null,"n":16,"split":"test","note":null},{"name":"Faithful sentences flagged (hard: unsupported or contradicted)","value":"0/64 (0)","unit":null,"n":64,"split":"test","note":null},{"name":"Key events flagged on faithful reports, missing or partly (missing only)","value":"9/68 (2/68)","unit":null,"n":68,"split":"test","note":null},{"name":"Plants found on ASR transcripts of the demo pair","value":"12/12","unit":null,"n":12,"split":"synthetic","note":"Two reports checked against MOSS-Transcribe-Diarize transcripts of synthetic audio, re-voiced 26 Sep 2026 with Decosa house voices (Kokoro-82M) and re-run: 2/19 faithful sentences flagged, one from a misheard name (first build: 12/12, 1/19)."}],"dataset":"12 invented incidents (5 police, 3 security, 2 EMS, 2 workplace), each with a timestamped transcript, a faithful report and a tampered report with six planted discrepancies (2 contradictions, 2 additions, 2 omissions). Dev 4 cases, test 8 cases run once after the prompts were frozen.","held_out":true,"caveats":["The building agent wrote the scenarios and the plants, so they are clear-cut; real reports are messier.","Synthetic incidents only: not measured on real incident reports against real body-worn-camera audio.","One run of the test split.","The check trusts the transcript: a mishearing the writer turns into a fact is marked as in the recording.","The first-draft writer's quality and the lite and best tiers are not measured."],"date":"2026-09-25","doc_url":"https://decosa.ai/metrics/evals/report-integrity"},"stack":{"summary":"For defence counsel and investigators, oversight boards, and agencies or employers that draft reports with AI. Give it the transcript of a body-worn-camera, security, EMS or workplace recording and the written report. Each sentence comes back as in the recording, partly, not in the recording, or the recording says otherwise, with the lines and times it rests on; key events on the recording (force, injuries, rights, consent, custody, what each person said) are looked for in the report. Draft mode writes the first draft from the recording, signs it, and seals a record of which sentences a person changed before signing, with the disclosure line California and Utah require. Real reports and footage are personal and often confidential, so the product runs on your own GPU; the hosted demo takes synthetic material only.","tagline":"Every sentence of an incident report checked against the recording it describes, with the times, plus the key events it leaves out.","deployment":"self-host-first","regulatory_note":"Checked 25 Sep 2026 against the enacted texts. California SB 524 (Stats. 2025, ch. 587; Penal Code 13663, in force 1 Jan 2026): an official report written fully or partly with AI must say so on each page or in its body and name the AI program, carry the officer's signature verifying the facts are true and correct, keep the first draft as long as the official report, and keep an audit trail of who used AI and the footage used; drafts other than the official report are not the officer's statement. Utah SB 180 (2025; Utah Code 53-25-601 and 53-25-602, in force 7 May 2025): a report created wholly or partly with generative AI must carry a disclaimer and the author's certification that they read and reviewed it for accuracy, and each agency needs a written AI policy. The disclosure and certification text here follows those statutes but is not legal advice. 'Not in the recording' means the transcript does not contain it, not that it is false: a camera can miss what a person saw. Criminal justice information held by agencies falls under the FBI CJIS Security Policy; this stack has no CJIS attestation, which is one more reason it is self-host first.","components":[{"id":"llm","role":"Key events, first-draft writer, sentence judge (the grounding checker) and event coverage","name":"Qwen3.8-27B (NVIDIA NVFP4)","hf_repo":"nvidia/Qwen3.8-27B-NVFP4","license":"Apache-2.0","params":"27.8B","quant":"NVFP4 (MLP) + FP8 (attention/GDN), FP8 KV cache, MTP speculative decoding k=3","vram_gb":57,"memory_gb_estimate":null,"engine":"vLLM 0.29.0","receipt_coverage":"strong","in_hosted_demo":true,"tiers":["standard"],"alternative_to":null},{"id":"asr","role":"Recording to a timed, speaker-labelled transcript (POST /report/transcribe, self-host)","name":"MOSS-Transcribe-Diarize 0.9B","hf_repo":"OpenMOSS-Team/MOSS-Transcribe-Diarize","license":"Apache-2.0","params":"0.9B","quant":"BF16","vram_gb":null,"memory_gb_estimate":null,"engine":"transformers (trust_remote_code) + moss_transcribe_diarize package (decosa-api services/diarize)","receipt_coverage":"partial","in_hosted_demo":false,"tiers":["standard","best","lite"],"alternative_to":null},{"id":"llm-lite","role":"Lite tier: the same four steps on a 48 GB card","name":"Gemma 4 26B A4B (instruction-tuned)","hf_repo":"google/gemma-4-26B-A4B-it","license":"Apache-2.0 (model card also links the Gemma 4 licence page)","params":"25.2B","quant":"BF16 weights; FP8 at load time (vLLM --quantization fp8) to fit a 48 GB card","vram_gb":null,"memory_gb_estimate":null,"engine":"vLLM 0.29.0 (Gemma4ForConditionalGeneration is in its model registry)","receipt_coverage":"none","in_hosted_demo":false,"tiers":["lite"],"alternative_to":null},{"id":"llm-best","role":"Best tier: a larger judge for long, many-speaker incidents","name":"DeepSeek-V4-Flash (NVIDIA NVFP4)","hf_repo":"nvidia/DeepSeek-V4-Flash-NVFP4","license":"MIT","params":"284B","quant":"NVFP4 experts + FP8 (about 159-176 GB of weights)","vram_gb":192,"memory_gb_estimate":null,"engine":"vLLM B12X community build, TP2 on 2x RTX PRO 6000, MTP draft fixed by the kit's patches","receipt_coverage":"none","in_hosted_demo":false,"tiers":["best"],"alternative_to":null},{"id":"glm-wanted","role":"Second judge, from another family","name":"GLM-5.3-Flash","hf_repo":"zai-org/GLM-5.3-Flash","license":"MIT","params":"321B","quant":"NVFP4 on NVIDIA (nvidia/GLM-5.3-Flash-NVFP4, about 170-186 GB, unconfirmed); MLX 4-bit on a Mac (165 GB)","vram_gb":null,"memory_gb_estimate":170,"engine":"SGLang SM120 build, TP2 on 2x 96 GB (vLLM is broken on sm_120 for this model, and the SGLang build hung on our server), or mlx-lm on a Mac with 192 GB or more","receipt_coverage":"none","in_hosted_demo":false,"tiers":["wanted"],"alternative_to":null}],"tiers":[{"id":"lite","label":"Lite · one 48 GB card","summary":"The same pipeline on a smaller mixture-of-experts model. Faster and cheaper; accuracy on this task unknown.","components":["llm-lite","asr"],"hardware":"1x L40S or RTX 6000 Ada 48 GB (not measured)","quality_evidence":[{"metric":"planted discrepancies found / faithful sentences flagged","value":"not measured yet","source":null}],"latency_note":"not measured yet","in_hosted_demo":false,"receipt_coverage":"partial","receipt_note":"Direct route: calls are attested by the box's key; no gateway receipts.","hosting":null},{"id":"standard","label":"Standard · the hosted demo, one 96 GB card","summary":"Qwen3.8-27B lists key events, drafts, judges every sentence and every event; MOSS-Transcribe-Diarize turns recordings into transcripts. Every model call receipted.","components":["llm","asr"],"hardware":"1x RTX PRO 6000 Blackwell 96 GB","quality_evidence":[{"metric":"held-out test, 8 synthetic incidents: planted additions / contradictions / omissions found","value":"16/16 / 16/16 / 16/16","source":"decosa-api docs/evals/report-integrity.md, measured on our server 2026-09-25, gateway route, one run; data written by the building agent, prompts frozen on a separate 4-case dev split"},{"metric":"faithful reports: sentences flagged / key events flagged as left out","value":"0/64 / 9/68 (2/68 as missing, 7 as partly)","source":"decosa-api docs/evals/report-integrity.md, measured on our server 2026-09-25, gateway route"},{"metric":"the two demo incidents on real speech-recognition transcripts","value":"4/4 additions, 4/4 contradictions, 4/4 omissions found; faithful sentences flagged 2/19 (1 partial, 1 contradicted)","source":"decosa-api docs/evals/report-integrity.md (asr run), measured on our server 2026-09-26, gateway route: the traffic-stop and bar-fight reports checked against MOSS-Transcribe-Diarize transcripts of their synthetic audio, re-voiced 26 Sep 2026 from macOS voices to Decosa house voices (Kokoro-82M) (the first build: 1/19 flagged, partial)"},{"metric":"real reports against real body-worn-camera audio, reviewed by a lawyer","value":"not measured yet","source":null}],"latency_note":"measured on our server: seconds to half a minute per check on a quiet gateway, up to a few minutes while the shared gateway was busy","in_hosted_demo":true,"receipt_coverage":"strong","receipt_note":"Hosted: gateway-signed receipt per model call. Self-hosted: attested by the box's key.","hosting":null},{"id":"best","label":"Best · DeepSeek-V4-Flash on two more cards","summary":"A larger judge for long, many-speaker incidents.","components":["llm-best","asr"],"hardware":"2x RTX PRO 6000 96 GB","quality_evidence":[{"metric":"planted discrepancies found / faithful sentences flagged","value":"not measured yet","source":null}],"latency_note":"not measured yet","in_hosted_demo":false,"receipt_coverage":"partial","receipt_note":"Not a hosted model: calls are attested by the box's key only.","hosting":null},{"id":"wanted","label":"Wanted · two large judges from different families","summary":"DeepSeek-V4-Flash and GLM-5.3-Flash each read the file, and a finding stands when they agree; disagreements go to the reviewer. Reports and transcripts stay on your own hardware, never on community providers. Not served yet.","components":["llm-best","asr","glm-wanted"],"hardware":"Your own hardware: 2x 96 GB cards for DeepSeek-V4-Flash plus 2x 96 GB for GLM-5.3-Flash, or one Mac Studio with 512 GB holding both 4-bit builds (156 + 165 GB, sizes from our Mac; not run together yet). Estimate.","quality_evidence":[{"metric":"planted discrepancies found / faithful sentences flagged","value":"not measured yet","source":null}],"latency_note":"not measured yet","in_hosted_demo":false,"receipt_coverage":"none","receipt_note":"On your own hardware its calls are attested by the box's key only: not a hosted model there, so no gateway receipts. Never sent to community providers.","hosting":"own-hardware"}],"alternates":[],"services":[{"name":"decosa-api","port":8445,"image":"${DECOSA_REGISTRY}/decosa-api:0.1.0","purpose":"Transcript intake, the check and draft runs, finalize, signing and the HTTP API (/report/*). No GPU. Binds 127.0.0.1 by default."},{"name":"decosa-llm","port":8000,"image":"${DECOSA_REGISTRY}/decosa-llm:0.1.0","purpose":"vLLM OpenAI endpoint for Qwen3.8-27B. Internal to the compose network."},{"name":"decosa-diarize","port":8092,"image":null,"purpose":"Optional, self-host only: MOSS-Transcribe-Diarize for transcripts made from recordings. No published image yet; built from services/diarize."}],"tools":[{"name":"decosa grounding checker (vertical 17)","url":"https://decosa.ai/apps/grounding","license":"AGPL-3.0-or-later (decosa-api)","purpose":"Per-sentence verdicts with evidence spans and a signed report; imported, not copied."},{"name":"cryptography (Python)","url":"https://cryptography.io/","license":"Apache-2.0 OR BSD-3-Clause","purpose":"Ed25519 signatures on the check, the first draft and the record."},{"name":"difflib (Python standard library)","url":"https://docs.python.org/3/library/difflib.html","license":"PSF-2.0","purpose":"Sentence-level diffs between the first draft and each saved version."}],"hardware":[{"tier":"1x RTX PRO 6000 Blackwell 96 GB","fits":true,"notes":"Measured: the hosted demo's Qwen3.8-27B runs on one of these cards on our server; the diarizer runs on the other card there."},{"tier":"1x L40S / RTX 6000 Ada 48 GB","fits":null,"notes":"Not measured. FP8 Qwen3.8-27B with a shorter context, or Gemma 4 26B A4B (lite)."},{"tier":"CPU only","fits":true,"notes":"Transcript parsing, finalize (diffs, provenance, the signed record) and verification need no GPU; the check itself needs the model."}],"latency":[{"lane":"check with key events, self-hosted (direct route, own vLLM)","typical_ms":6000,"source":"measured on our server 2026-09-25, self-host sandbox against the local Qwen3.8-27B vLLM; 13 s for a draft with its check"},{"lane":"check of a 10-12 sentence report with key events, quiet gateway","typical_ms":15000,"source":"decosa-api docs/evals/report-integrity.md, dev runs on our server 2026-09-25 (8-36 s), gateway route"},{"lane":"check of an 8-11 sentence report with key events, busy shared gateway","typical_ms":83000,"source":"decosa-api docs/evals/report-integrity.md, test runs on our server 2026-09-25 (43-168 s, mean 83 s), gateway shared with other evaluation jobs"},{"lane":"transcribe a 100 s recording (self-host)","typical_ms":5900,"source":"measured on our server 2026-09-25, MOSS-Transcribe-Diarize on an RTX PRO 6000"},{"lane":"finalize (diff, provenance, signed record)","typical_ms":50,"source":"estimate: no model call"}],"benchmark":null,"notes":[]},"buyer_facts":[{"label":"What it checks","value":"A written report against the transcript of its recording: typed, WebVTT or SRT, or made from the audio by the diarizer on your box."},{"label":"Data kept","value":"Nothing. Transcripts, reports and records live in memory for the request; logs carry counts and timings only. You keep the signed statement and record."},{"label":"What leaves the box (hosted demo)","value":"The transcript and report go to Qwen3.8-27B through our gateway; the gateway's receipts hold hashes, not text. Self-hosted: nothing leaves."},{"label":"Model calls per check","value":"One per report sentence, one for the key-event list, one per key event: about 9 for a 10-sentence report (284 calls over 32 eval runs)."},{"label":"Typical run cost","value":"A fraction of a cent per check at the open-market median price for Qwen3.8-27B. Each run shows its own measured cost."},{"label":"Disclosure wording","value":"California SB 524 and Utah SB 180 text built in, plus a neutral option; the signed record holds the disclosure and the author's certification."}],"data_handling":{"page":"/data#report-integrity","self_host":{"level":"confidential","leaves":"nothing","summary":"Runs on your machine; nothing is sent to Decosa or a third party by default."},"hosted":{"level":"operator-processed","demo_only":true,"summary":"Hosted demo on sample or public data only; self-host for real data.","gpus":"operator-contracted","third_parties":[],"retention":"Nothing. Transcripts, reports and records live in memory for the request; logs carry counts and timings only. You keep the signed statement and record.","used_for_training":false,"encrypted_while_processed":false},"sealed_tier":{"applies":false,"note":"The sealed tier (raw chat only, never use-case pipelines) is paused at launch (/docs/sealed-tier)."},"external_calls":[]},"console":{"href":"/legal/report-integrity","input":"report","lanes":[{"id":"sentences","title":"Report vs recording","kind":"list"},{"id":"events","title":"Key events left out","kind":"list"},{"id":"record","title":"Signed record","kind":"json"}],"samples":[{"n":1,"id":"traffic-stop-audio","title":"Traffic stop audio","deep_link":"/legal/report-integrity?sample=1&autorun=0"},{"n":2,"id":"bar-fight-audio","title":"Bar fight audio","deep_link":"/legal/report-integrity?sample=2&autorun=0"},{"n":3,"id":"retail-detention","title":"Retail detention","deep_link":"/legal/report-integrity?sample=3&autorun=0"},{"n":4,"id":"ems-fall","title":"Ems fall","deep_link":"/legal/report-integrity?sample=4&autorun=0"},{"n":5,"id":"forklift","title":"Forklift","deep_link":"/legal/report-integrity?sample=5&autorun=0"},{"n":6,"id":"traffic-stop-faithful","title":"Traffic stop faithful","deep_link":"/legal/report-integrity?sample=6&autorun=0"}],"deep_link_params":{"sample":"1-based index into samples, or a sample id","autorun":"1 = start the run once the sample is loaded; 0 (default) = only preselect","reduce-motion":"1 = turn off animations"}},"api":{"base":"https://api.decosa.ai","contract":"/api/contract.json","contract_markdown":"/api/contract.md","reference":"/docs/api","keys":"/account/keys"},"prompts":{"hosted":"/prompts/report-integrity-hosted.md","selfhost":"/prompts/report-integrity-selfhost.md","assemble":"/prompts/report-integrity-assemble.md","mac":"/prompts/report-integrity-mac.md"},"rehearsal":{"bundle":"/samples/report-integrity.zip","bundle_url":"https://decosa.ai/samples/report-integrity.zip","folder":"/samples/report-integrity/","expected":"/samples/report-integrity/expected.json","files":["/samples/report-integrity/expected.json","/samples/report-integrity/inputs/forklift-interview-transcript.txt","/samples/report-integrity/inputs/incident-report.txt","/samples/report-integrity/inputs/what-was-planted.json"],"bytes":4050,"checks":["every sentence got a verdict (none errored)","'thirty feet' (the recording says ten) is flagged","'the backup alarm was not working' (the recording says it was) is flagged","'looking at his phone' (never said on the recording) is flagged","'two prior forklift incidents' (never said on the recording) is flagged","the faithful 'walking speed' sentence is supported","the left-out events (Kevin's shoulder, the declined clinic) are listed as missing","the statement was signed by this server","the signature is valid","the statement matches the report text","an edited report no longer matches the signed statement","a statement with one verdict changed no longer verifies","every model call has a signed receipt"],"licence":"Synthetic: a typed transcript and report written for Decosa (invented people, places and vehicles). Part of decosa-api, AGPL-3.0-or-later.","about":"A synthetic forklift-incident interview (timestamped transcript) and a supervisor's report written with planted problems: two statements the recording contradicts (thirty feet instead of ten; the backup alarm 'not working'), two statements the recording never makes (looking at a phone, prior incidents), and two events left out (the box hitting Kevin's shoulder, Kevin declining the clinic). Each planted sentence must be flagged, the faithful sentences supported, the left-out events listed as missing, and the signed statement must verify and catch an edit.","run":{"containers":"docker compose exec api python scripts/rehearse.py report-integrity","checkout":"python scripts/rehearse.py report-integrity --bundle report-integrity.zip --base-url http://127.0.0.1:8445","mac":".venv/bin/python scripts/rehearse.py report-integrity"},"guidance":"Set up with a coding agent (we recommend Claude Code with Claude Opus 5.5; any capable coding agent works) on mock data only, run the rehearsal until every check passes, then run your own data locally yourself. Never give the agent real data during setup."},"hardware_fit":{"check":"/self-host/hardware?use=report-integrity","data":"/api/hardware.json","tiers":[{"id":"lite","gpu_gb":44,"basis":"estimate","unknown":[]},{"id":"standard","gpu_gb":61.6,"basis":"estimate","unknown":[]},{"id":"best","gpu_gb":196,"basis":"estimate","unknown":[]},{"id":"wanted","gpu_gb":388,"basis":"estimate","unknown":[]}],"mac":{"fit":"full","memory_gb":32}},"links":{"page":"/legal/report-integrity","json":"/use-cases/report-integrity.json","metrics":"/metrics/report-integrity","console":"/legal/report-integrity","console_sample":"/legal/report-integrity?sample=1&autorun=0","stack":"/legal/report-integrity#stack","try_live":"/legal/report-integrity","watch":"/legal/report-integrity","build":"/legal/report-integrity#build","self_host":"/legal/report-integrity#self-host","prompts":{"hosted":"/prompts/report-integrity-hosted.md","selfhost":"/prompts/report-integrity-selfhost.md","assemble":"/prompts/report-integrity-assemble.md","mac":"/prompts/report-integrity-mac.md"}}}