Skip to content
decosa
LiveHostedSelf-hostMacSelf-host first for real data

Check a police report against bodycam

A sentence-by-sentence check of the report against the recording, with the times to cue up, plus the key events the report leaves out.

Measured16/16 (16/16)Planted contradictions with the recording found (found exactly)
On production97 smedian on production (2026-09-25); slower when the service is busy
List price~$0.42 per 100 reportsmeasured, at list price

Built on: Speaker diarization, Grounding, Signed record

Loading the tool…

Use it your way

Use it from your codeThe hosted API with your key, and prompts to paste into a coding agent
Self-host · your GPUs · recommended

Run it yourself, on request

  • The same open models and app, on 1× RTX PRO 6000 (96 GB) for Qwen3.8-27B and the diarizer; the check itself needs no GPU.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Hosted · by Decosa

Get an API key

  • Call the report integrity API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Synthetic, public or test data only: real confidential data belongs on your own hardware.

Build with it

Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.

Base URL
https://api.decosa.ai
Auth
Authorization: Bearer $DECOSA_API_KEY (or a demo session token)
Tool id
report-integrity

Use the hosted API

# Decosa report integrity: use the hosted API

You are wiring Decosa's report-integrity check into this project. It takes the transcript of a recording (body-worn
camera, security camera, EMS crew, workplace interview) and a written incident report, and says for each sentence of the
report whether the recording supports it: `supported` (in the recording), `partial`, `unsupported` (not in the
recording) or `contradicted` (the recording says otherwise), with the transcript lines and times it rests on. It also
lists key events on the recording (force, injuries, rights, consent, custody, what each person said) and whether the
report covers them. Every model call has its own signed receipt, and the result is signed by the server over hashes
only. Draft mode writes a first draft from the transcript, signs it, and `/report/finalize` seals a record of which
sentences a person changed before signing. Use only what is listed below. If you need something else, stop and ask me.

- Base URL: `https://api.decosa.ai`
- Health check: `GET https://api.decosa.ai/healthz`.
- **Hosted use is for synthetic or public material only.** Real incident reports, footage and discovery belong on a
  self-hosted box (the self-host prompt). Say so wherever this is wired in.
- "Not in the recording" means the transcript does not contain it, not that it is false. Show that wording with results.

## Auth: API key (or a demo session)
1. Preferred: an API key (`dk_…`) from "Get an API key" on the tool page. Keep it in an environment variable,
   `DECOSA_API_KEY`, never in code. Send `Authorization: Bearer $DECOSA_API_KEY`.
2. Without a key: `POST https://api.decosa.ai/demo/session` with `{"vertical": "report-integrity"}` returns `{"token", "expires_at", "budget"}`.
   a limited number of sessions per network per hour (the current limits are in `demo_sessions` of GET /healthz); each session has a token allowance (its `budget`) (about three checks of a 12-sentence report).
   Over a limit: HTTP 429 with `Retry-After`. One run at a time per demo token (409 otherwise).
3. A check needs about 160 generated tokens per report sentence plus about 1,700 for the key events (402 otherwise).

## Endpoints
- `POST /report/check` (token). Body: `{"report": "...", "transcript": "[00:21] OFFICER REYES: ...", "author"?: "Officer Reyes", "kind"?: "police|security|ems|workplace|other", "speakers"?: {"S01": "Officer Reyes"}, "events"?: true}`.
  - Give exactly one of `transcript` (text: one utterance per line starting with a time like `[01:23]`, or WebVTT / SRT),
    `segments` (`[{start, end, speaker, text}]`, seconds, as a diarizer returns them) or `sample_id` (from `/report/samples`;
    with a sample, `"report": {"sample": "tampered"}` or `{"sample": "faithful"}` uses its bundled report).
  - Limits: report up to 12,000 characters and 60 sentences; transcript up to 800 lines and 80,000 characters.
  - `author` is who "I" is in the report. `"events": false` skips the key-event pass (cheaper).
  - JSON by default: `{sentences: [{i, text, verdict, label, certainty, confidence, reason, evidence: [{line, at, t_ms, speaker, quote, role}], receipt_ids}], events, omissions: [{id, what, lines, at, category, verdict: "covered"|"partly"|"missing", sentences, reason}], counts, omission_counts, flagged, missing, signed, exhibit_md, receipts, transcript, budget, note}`.
  - With `Accept: text/event-stream` (or `"stream": true`): `ready`, `events`, `sentences`, then `receipt` + `sentence`
    and `receipt` + `omission` events as they finish (not in order), then `report`, `budget`, `done`.
- `POST /report/draft` (token): same transcript fields, plus `author` and `kind`. Returns the check fields plus
  `draft: {text, sentences: [{i, text, lines, at}], to_confirm}` and `first_draft` (a signed statement: draft hash,
  sentence hashes, transcript hash, writer prompt hash, receipts).
- `POST /report/finalize` (token, no model call): `{"first_draft": {...}, "draft": "<the draft text as returned>", "versions": [{"text": "...", "author"?: "...", "note"?: "..."}], "jurisdiction": "CA"|"UT"|"none", "certification": {"name": "...", "accepted": true}}`
  → `{provenance: [{i, text, source: "ai"|"ai_edited"|"human", from_draft}], edits, counts, share_pct, disclosure: {text, placement, law, keep}, certification, record, self_check}`.
  The last version is the report being signed. `record` is a signed hash chain; check it with `POST /record/verify {"record": ...}`.
- `POST /report/verify` (no token) `{"statement": {...}, "report"?: "...", "transcript_sha256"?: "..."}` → `{valid_signature, signed_by_this_server, text_matches?, transcript_matches?}`.
- `GET /report/info`, `GET /report/samples`, `GET /attest/signing-key` (no token). `POST /report/transcribe` is self-host only (503 here).

## Example: check a report and print what needs a person's eyes (Python, `pip install httpx`)
```python
import httpx, os
API = "https://api.decosa.ai"
H = {"Authorization": f"Bearer {os.environ['DECOSA_API_KEY']}"}
body = {"transcript": open("transcript.txt").read(), "report": open("report.txt").read(), "author": "Officer Reyes", "kind": "police"}
r = httpx.post(f"{API}/report/check", json=body, headers=H, timeout=300)
r.raise_for_status()
js = r.json()
for s in js["sentences"]:
    if s["verdict"] in ("partial", "unsupported", "contradicted"):
        where = "; ".join(f"{e['at']} {e['quote']}" for e in s["evidence"][:2])
        print(f"{s['label']}: {s['text']}\n   {s['reason']}\n   {where}")
for o in js["omissions"]:
    if o["verdict"] != "covered":
        print(f"{o['verdict']}: {o['what']} ({', '.join(o['at'])})")
open("check.json", "w").write(__import__("json").dumps(js["signed"]))
open("exhibit.md", "w").write(js["exhibit_md"])
```

## Verify a statement yourself (`pip install cryptography`)
```python
import hashlib, json, urllib.request
from cryptography.hazmat.primitives.asymmetric.ed25519 import Ed25519PublicKey
pub = json.load(urllib.request.urlopen("https://api.decosa.ai/attest/signing-key"))["pubkey"]
st = json.load(open("check.json"))
assert st["signer"] == pub
body = {k: v for k, v in st.items() if k != "sig"}
msg = st["v"].encode() + b"\n" + json.dumps(body, sort_keys=True, separators=(",", ":"), ensure_ascii=False).encode()
Ed25519PublicKey.from_public_bytes(bytes.fromhex(pub)).verify(bytes.fromhex(st["sig"]), msg)
assert hashlib.sha256(open("report.txt").read().encode()).hexdigest() == st["report"]["sha256"]
```

Run it yourself (containers)

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

# Decosa report integrity: run it yourself (containers)

You are setting up Decosa report integrity on this machine, so incident reports, recordings and discovery never leave
it. It checks each sentence of a report against the transcript of its recording, lists key events the report leaves
out, drafts reports with a signed first draft, and seals a record of edits. Nothing is sent to Decosa's hosted API.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/report-integrity.zip (4 KB, 13 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py report-integrity` (the api image carries the same bundle under /app/rehearsal/report-integrity/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py report-integrity --bundle report-integrity.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "every sentence got a verdict (none errored)", "'thirty feet' (the recording says ten) is flagged", "'the backup alarm was not working' (the recording says it was) is flagged"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
   instructions for this distribution (docs.docker.com/engine/install). Install the NVIDIA container toolkit and check
   `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
   `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
   Read it. Keep the `llm` service (Qwen3.8-27B on vLLM) and the `api` service. For the `api` service set
   `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b`, keep its data on a
   named volume, and bind every port to 127.0.0.1. For recordings, also keep the `diarize` service and set
   `DECOSA_DIARIZE_URL=http://diarize:8092` and `DECOSA_REPORT_AUDIO=1` on the api.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check (the first start
   downloads about 20 GB of weights).
4. Check: `curl -fsS http://127.0.0.1:<PORT>/report/info` lists the verdicts and the law notes; `GET /attest/signing-key`
   shows this box's public key. Show me the key: it is what others pin to verify my checks.
5. Smoke test: get a token with `POST /demo/session {"vertical":"report-integrity"}`, then
   `POST /report/check {"sample_id": "traffic-stop-audio", "report": {"sample": "tampered"}}`. Expect the 62 mph sentence
   `contradicted`, the odor sentence `unsupported`, the refused search among the `missing` key events, and every receipt
   `attested`. Then `POST /report/verify {"statement": <signed>}`: `valid_signature` and `signed_by_this_server` true.
6. Report back: the public key and key id, the smoke-test verdicts, and how long the check took.

"Not in the recording" means the transcript does not contain it, not that it is false. This is not legal advice.

Off by default. Joining as a provider serves other people's requests on this GPU; never do it on a box that holds
incident reports or discovery. If I ask for it later, follow the Provide page instead of improvising.

No NVIDIA GPU? This tool also runs entirely on an Apple Silicon Mac (MLX, 32 GB of unified memory or more): use https://decosa.ai/prompts/report-integrity-mac.md instead.
Run it on your own hardwareWhat it needs, and the prompt that sets it up

Run it on your own GPU

Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.

  • CPU only, 64 GB RAMDoesn't fit

    Qwen3.8-27B (NVIDIA NVFP4) needs a GPU.

  • GeForce RTX 4090standard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)

  • GeForce RTX 5090standard tierRuns

    The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

  • 2x GeForce RTX 5090standard tierRuns

    The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

  • L40Sstandard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.

  • H100 80 GB (SXM)standard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.

  • RTX PRO 6000 Blackwell 96 GBstandard tierRuns

    The standard tier fits (61.6 of 96 GB).

  • 2x RTX PRO 6000 Blackwell 96 GBbest tierRuns

    The standard tier fits (61.6 of 192 GB). The best tier fits too.

  • Apple M3 Ultra (Mac Studio), 96 GBstandard tierRuns

    The standard tier fits (32 of 96 GB).

  • Apple M5 Max, 64 GBstandard tierRuns

    The standard tier fits (32 of 64 GB).

Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

  1. 1

    Check the GPU, Docker and the NVIDIA Container Toolkit

    The driver must see the GPU, and Docker must be able to pass it into a container.

    nvidia-smi
    docker compose version
    docker run --rm --gpus all ubuntu nvidia-smi
  2. 2

    Fetch the compose file

    One file describes the API, the speech model and the language model as services.

    mkdir -p ~/decosa && cd ~/decosa
    curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml
  3. 3

    Pull and start

    The first start downloads pinned model weights, tens of gigabytes.

    docker compose pull
    docker compose up -d
  4. 4

    Check health

    Wait until the API reports ok with both models loaded. Then point your app at the local base URL.

    curl -fsS http://localhost:<PORT>/healthz
    # {"ok": true, "asr": true, "llm": true, ...}
    curl -fsS -X POST http://localhost:<PORT>/demo/session \
      -H 'Content-Type: application/json' -d '{"vertical":"report-integrity"}'

Set up with a coding agent, rehearse on mock data, then go private

  1. Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
  2. Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's expected.json. Every check must print PASS.
  3. Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
Rehearsal command
docker compose exec api python scripts/rehearse.py report-integrity

Download the mock-data bundle (4 KB, 13 checks)expected.json

A synthetic forklift-incident interview (timestamped transcript) and a supervisor's report written with planted problems: two statements the recording contradicts (thirty feet instead of ten; the backup alarm 'not working'), two statements the recording never makes (looking at a phone, prior incidents), and two events left out (the box hitting Kevin's shoulder, Kevin declining the clinic). Each planted sentence must be flagged, the faithful sentences supported, the left-out events listed as missing, and the signed statement must verify and catch an edit.

What the rehearsal checks
  • every sentence got a verdict (none errored)
  • 'thirty feet' (the recording says ten) is flagged
  • 'the backup alarm was not working' (the recording says it was) is flagged
  • 'looking at his phone' (never said on the recording) is flagged
  • 'two prior forklift incidents' (never said on the recording) is flagged
  • the faithful 'walking speed' sentence is supported
  • the left-out events (Kevin's shoulder, the declined clinic) are listed as missing
  • the statement was signed by this server
  • the signature is valid
  • the statement matches the report text
  • an edited report no longer matches the signed statement
  • a statement with one verdict changed no longer verifies
  • every model call has a signed receipt

Licence: Synthetic: a typed transcript and report written for Decosa (invented people, places and vehicles). Part of decosa-api, AGPL-3.0-or-later.

Prompt for your coding agent

# Decosa report integrity: run it yourself (containers)

You are setting up Decosa report integrity on this machine, so incident reports, recordings and discovery never leave
it. It checks each sentence of a report against the transcript of its recording, lists key events the report leaves
out, drafts reports with a signed first draft, and seals a record of edits. Nothing is sent to Decosa's hosted API.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/report-integrity.zip (4 KB, 13 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py report-integrity` (the api image carries the same bundle under /app/rehearsal/report-integrity/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py report-integrity --bundle report-integrity.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "every sentence got a verdict (none errored)", "'thirty feet' (the recording says ten) is flagged", "'the backup alarm was not working' (the recording says it was) is flagged"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
   instructions for this distribution (docs.docker.com/engine/install). Install the NVIDIA container toolkit and check
   `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
   `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
   Read it. Keep the `llm` service (Qwen3.8-27B on vLLM) and the `api` service. For the `api` service set
   `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b`, keep its data on a
   named volume, and bind every port to 127.0.0.1. For recordings, also keep the `diarize` service and set
   `DECOSA_DIARIZE_URL=http://diarize:8092` and `DECOSA_REPORT_AUDIO=1` on the api.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check (the first start
   downloads about 20 GB of weights).
4. Check: `curl -fsS http://127.0.0.1:<PORT>/report/info` lists the verdicts and the law notes; `GET /attest/signing-key`
   shows this box's public key. Show me the key: it is what others pin to verify my checks.
5. Smoke test: get a token with `POST /demo/session {"vertical":"report-integrity"}`, then
   `POST /report/check {"sample_id": "traffic-stop-audio", "report": {"sample": "tampered"}}`. Expect the 62 mph sentence
   `contradicted`, the odor sentence `unsupported`, the refused search among the `missing` key events, and every receipt
   `attested`. Then `POST /report/verify {"statement": <signed>}`: `valid_signature` and `signed_by_this_server` true.
6. Report back: the public key and key id, the smoke-test verdicts, and how long the check took.

"Not in the recording" means the transcript does not contain it, not that it is false. This is not legal advice.

Off by default. Joining as a provider serves other people's requests on this GPU; never do it on a box that holds
incident reports or discovery. If I ask for it later, follow the Provide page instead of improvising.

No NVIDIA GPU? This tool also runs entirely on an Apple Silicon Mac (MLX, 32 GB of unified memory or more): use https://decosa.ai/prompts/report-integrity-mac.md instead.

Help me customise for my hardware

Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.

Hardware

GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page

RunsReport integrity on GeForce RTX 5090: use the Standard · the hosted demo, one 96 GB card tier

The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

Standard · the hosted demo, one 96 GB card: what changesuses estimates

  • Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
Memory per component
  • Key events, first-draft writer, sentence judge: Qwen3.8-27B (NVIDIA NVFP4). ~57.6 GB (at least ~28 GB), weights 21.4 GB (from stack.json). Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)
  • Recording to a timed, speaker-labelled transc...: MOSS-Transcribe-Diarize 0.9B. ~4 GB, weights 1.8 GB (estimate). MOSS-Transcribe-Diarize 0.9B: BF16 weights 1.8 GB (clinical stack.json). Working memory for long recordings is not measured; 4 GB is an estimate.

Expected speed

Not measured.

Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.

Setup prompt for this hardware

The self-host prompt for Report integrity, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.

# Set up Report integrity on my hardware

Fetch https://decosa.ai/prompts/report-integrity-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied.

## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=report-integrity)

Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4).
Quality tier: Standard · the hosted demo, one 96 GB card (standard). Fit check: runs with changes, about 32 GB of 32 GB used; some memory numbers are estimates, not measurements.

First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything.

Use these components (the setup below describes the standard tier; change it to match):
- Key events, first-draft writer, sentence judge: Qwen3.8-27B (NVIDIA NVFP4) (nvidia/Qwen3.8-27B-NVFP4), 57.6 GB. Change: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
- Recording to a timed, speaker-labelled transc...: MOSS-Transcribe-Diarize 0.9B (OpenMOSS-Team/MOSS-Transcribe-Diarize), 4 GB

GPU placement (set each service's device and its vLLM --gpu-memory-utilization to about the share shown):
- GPU 0: Qwen3.8-27B (NVIDIA NVFP4) ~28 GB (88%), MOSS-Transcribe-Diarize 0.9B ~4 GB (13%); about 0 GB left

During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed.

The stack's own component list and compose layout: https://decosa.ai/prompts/report-integrity-assemble.md

Or on a Mac Studio

No NVIDIA GPU needed: every model this tool uses runs natively on Apple Silicon through MLX. Any M-series Mac with 32 GB of unified memory or more. Measured speeds and what runs where

From a checkout of decosa-api, one command sets up the models and the API: scripts/mac/setup.sh

Mac prompt for your coding agent

# Decosa Report integrity: run it on this Mac (Apple Silicon, no NVIDIA GPU)

You are setting up the Decosa Report integrity on this Mac, natively on Apple Silicon. The models run on the Mac's GPU
through MLX and decosa-api runs from a git checkout with `uv`. Docker is not used for the models, because Docker on
macOS cannot reach the GPU. Nothing is sent to Decosa's hosted API.

Every model this tool needs runs on the Mac. It needs 32 GB of unified memory or more.

Ask me before any command that needs sudo or installs software with Homebrew, and show me the command first. Never stop
or kill a process this setup did not start; if a port is taken, pick another one.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/report-integrity.zip (4 KB, 13 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `.venv/bin/python scripts/rehearse.py report-integrity` in the decosa-api checkout (the key comes from ~/.decosa-mac/api.key).
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "every sentence got a verdict (none errored)", "'thirty feet' (the recording says ten) is flagged", "'the backup alarm was not working' (the recording says it was) is flagged"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## What runs where

| Part | On an NVIDIA GPU | On this Mac | Status |
|---|---|---|---|
| Key events, first-draft writer, sentence judge (the grounding checker) and event coverage | NVFP4 on vLLM 0.29 (Blackwell) | MLX 4-bit (EigenLabs/Qwen3.8-27B-4bit) on mlx_lm.server 0.31.3; oMLX 0.6.1 with MTP as an option | Runs, measured |
| Recording to a timed, speaker-labelled transcript (POST /report/transcribe, self-host) | transformers on CUDA | MLX 8-bit (vanch007/mlx-MOSS-Transcribe-Diarize-8bit) on mlx-audio, scripts/mac/diarize_server.py | Runs, measured |

## Steps
1. Check the machine: `uname -m` must print `arm64` (an M-series chip; Intel Macs cannot run MLX), and
   `sysctl -n hw.memsize` should be at least 32 GB for this tool. Check about 30 GB of free disk with
   `df -h ~`. Show me the chip (`sysctl -n machdep.cpu.brand_string`) and the memory.
2. Tools: `uv --version`. If it is missing, ask me, then `brew install uv`.
3. Code: `git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host> ~/decosa-api` (access required) and `cd ~/decosa-api`.
   Check that `scripts/mac/setup.sh` exists; if it does not, the checkout is too old: stop and tell me.
4. Start everything with one command: `scripts/mac/setup.sh`. It creates `.venv` (decosa-api)
   and `.venv-mac` (MLX, mlx-lm, mlx-audio), downloads the weights with the Hugging Face CLI (about 16 GB for the
   language model), starts the model servers and decosa-api on 127.0.0.1, and mints a local API key
   into `~/.decosa-mac/api.key` (mode 0600). The first run takes a while because of the downloads; later runs reuse them.
   If a download fails with 401 or 403, ask me for a Hugging Face token and set `HF_TOKEN`.
5. Check health: `scripts/mac/setup.sh status` shows each server, and `curl -fsS http://127.0.0.1:8445/healthz` must
   report `"llm": true`. `curl -fsS http://127.0.0.1:8445/attest/signing-key` shows this Mac's public key:
   show it to me, because it is what others pin to check the receipts and records this Mac signs.
6. Smoke test: `.venv/bin/python scripts/mac/bench_usecases.py report-integrity`. It runs the tool's own sample end to end
   against the local API with the local key and prints `ok`, the wall time, the model calls and the receipts.
   `ok=True` is the pass condition. If it fails, read `~/.decosa-mac/logs/*.log` and tell me what you found.
7. Point the app at it: the API is `http://127.0.0.1:8445` with `Authorization: Bearer $(cat ~/.decosa-mac/api.key)`,
   the same routes as the hosted API. To stop everything: `scripts/mac/setup.sh stop`.
8. Report back: the chip and memory, the public key, the smoke-test result and its time, and the output of
   `scripts/mac/setup.sh status`.

## Good to know
- Receipts: every model call is signed with this Mac's own Ed25519 key and names the exact MLX weights
  (`qwen3.8-27b-mlx-4bit` with a hash of the downloaded files). There is no gateway countersignature on a
  self-hosted Mac.
- The weights are a 4-bit MLX build of the same open models, not the NVFP4 build the hosted route and the published
  evals use. Expect small differences in wording and scores.
- Faster drafting: `scripts/mac/setup.sh stop && scripts/mac/setup.sh --engine omlx` serves the
  model with oMLX and multi-token prediction (about 2x faster for a single long answer, no faster for many parallel
  calls; typed judgments then use sampling because oMLX returns no log-probabilities).
- Measured speeds for a Mac Studio M3 Ultra and the memory each tool needs: https://decosa.ai/mac. Full details:
  `docs/self-host-mac.md` in the checkout.

The proof

How we tested itEval results and end-to-end checks, hosted and self-hosted, with dates

Verified end to end

Hosted: verified 25 Sep 2026 · measured 25 Sep 2026: · p50 97 s · ~$0.004 per run · 21 receipts

Loading the nightly status…

Self-host: verified 25 Sep 2026 · fresh clone of decosa-api into a clean directory on our server, api image built from docker/api/Dockerfile, compose api service with a named data volume, direct route, local signing; torn down after

Measured cost to run: about $0.42 per 100 reports (hosted, 25 Sep 2026, partly estimated). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.

Verified on 2026-09-25: the image builds, the service starts, and both smoke tests in the assembly prompt pass end to end against a local model server equivalent to the documented one (the already-running Qwen3.8-27B vLLM on 127.0.0.1:8114, reached with network_mode host instead of the compose llm service); model-server startup itself not re-verified, and the diarizer path was not run. Check: 2 contradicted, 2 unsupported, refused search and the one-beer answer missing, 20 attested receipts, signature verified, 6 s. Draft: 12 sentences with two likely mishearings listed for the author, finalize 12 AI / 1 human, record verified, 13 s.

Known limits (7)
  • Hosted verification ran on the pre-release server (decosa-api the pre-release branch on our server, gateway route) driven from the branch site in headless Chromium, including 390 px; the production API gets this vertical when the branch merges.
  • The check trusts the transcript. Speech-recognition errors pass through: in the demo the diarizer heard 'Camera on' as 'Cameron', and the draft named a bartender Cameron; the check marks it 'in the recording'. The writer now lists likely mishearings for the author to confirm.
  • 'Not in the recording' is not 'false': what a camera cannot hear (smells, what someone saw) is flagged and needs the author's own account.
  • Measured on 12 synthetic incidents written by the building agent, with clear-cut plants; real reports are messier. Not yet measured on real body-worn-camera audio or with a lawyer reviewing.
  • Key-event coverage is noisier than the sentence check: 9 of 68 events on faithful reports were marked partly or missing, mostly detail a reader would not miss.
  • Hosted runs took 80-130 s while the shared gateway was busy (8-36 s when quiet).
  • Audio intake (/report/transcribe) is self-host only.

Eval results, nightly checks and cost per runVerify a run

How it's builtThe steps, the models and what each one checks
Self-host · your GPUs · recommended

Run it yourself, on request

  • The same open models and app, on 1× RTX PRO 6000 (96 GB) for Qwen3.8-27B and the diarizer; the check itself needs no GPU.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Hosted · by Decosa

Get an API key

  • Call the report integrity API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Synthetic, public or test data only: real confidential data belongs on your own hardware.
The open stack

Every sentence of an incident report checked against the recording it describes, with the times, plus the key events it leaves out.

For defence counsel and investigators, oversight boards, and agencies or employers that draft reports with AI. Give it the transcript of a body-worn-camera, security, EMS or workplace recording and the written report. Each sentence comes back as in the recording, partly, not in the recording, or the recording says otherwise, with the lines and times it rests on; key events on the recording (force, injuries, rights, consent, custody, what each person said) are looked for in the report. Draft mode writes the first draft from the recording, signs it, and seals a record of which sentences a person changed before signing, with the disclosure line California and Utah require. Real reports and footage are personal and often confidential, so the product runs on your own GPU; the hosted demo takes synthetic material only.

Deployment
Self-host first
Regulatory
Checked 25 Sep 2026 against the enacted texts. California SB 524 (Stats. 2025, ch. 587; Penal Code 13663, in force 1 Jan 2026): an official report written fully or partly with AI must say so on each page or in its body and name the AI program, carry the officer's signature verifying the facts are true and correct, keep the first draft as long as the official report, and keep an audit trail of who used AI and the footage used; drafts other than the official report are not the officer's statement. Utah SB 180 (2025; Utah Code 53-25-601 and 53-25-602, in force 7 May 2025): a report created wholly or partly with generative AI must carry a disclaimer and the author's certification that they read and reviewed it for accuracy, and each agency needs a written AI policy. The disclosure and certification text here follows those statutes but is not legal advice. 'Not in the recording' means the transcript does not contain it, not that it is false: a camera can miss what a person saw. Criminal justice information held by agencies falls under the FBI CJIS Security Policy; this stack has no CJIS attestation, which is one more reason it is self-host first.
Architecture
Text description

A recording (self-host only) goes through MOSS-Transcribe-Diarize to a timed, speaker-labelled transcript; a typed, WebVTT or SRT transcript can be given instead. decosa-api numbers the lines with their times and hashes them. Qwen3.8-27B lists the key events on the recording, writes a cited first draft in draft mode, judges every report sentence against the transcript (the grounding checker) and looks for every key event in the report. Outputs: sentence verdicts with times, key events left out, a signed check, a signed first draft, and after a person edits and certifies, a signed hash-chained record of which sentences are AI, AI-edited or human, with the disclosure line. Hosted calls get gateway-signed receipts; a self-hosted box signs with its own key.

Architecture

At a glance

What it checks
A written report against the transcript of its recording: typed, WebVTT or SRT, or made from the audio by the diarizer on your box.
Data kept
Nothing. Transcripts, reports and records live in memory for the request; logs carry counts and timings only. You keep the signed statement and record.
What leaves the box (hosted demo)
The transcript and report go to Qwen3.8-27B through our gateway; the gateway's receipts hold hashes, not text. Self-hosted: nothing leaves.
Model calls per check
One per report sentence, one for the key-event list, one per key event: about 9 for a 10-sentence report (284 calls over 32 eval runs).
Typical run cost
A fraction of a cent per check at the open-market median price for Qwen3.8-27B. Each run shows its own measured cost.
Disclosure wording
California SB 524 and Utah SB 180 text built in, plus a neutral option; the signed record holds the disclosure and the author's certification.
Quality tiers

Pick the tier for the quality you need

Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.

  • Lite

    one 48 GB card

    The same pipeline on a smaller mixture-of-experts model. Faster and cheaper; accuracy on this task unknown.

    Models
    • MOSS-Transcribe-Diarize 0.9B
    • Gemma 4 26B A4B (instruction-tuned)
    Hardware
    1x L40S or RTX 6000 Ada 48 GB (not measured)
    Quality evidence
    • planted discrepancies found / faithful sentences flaggednot measured yet
    Latency
    not measured yet
    Verification
    Proof: partialSelf-host onlyDirect route: calls are attested by the box's key; no gateway receipts.
  • In the hosted demo

    Standard

    the hosted demo, one 96 GB card

    Qwen3.8-27B lists key events, drafts, judges every sentence and every event; MOSS-Transcribe-Diarize turns recordings into transcripts. Every model call receipted.

    Models
    • Qwen3.8-27B (NVIDIA NVFP4)
    • MOSS-Transcribe-Diarize 0.9B
    Hardware
    1x RTX PRO 6000 Blackwell 96 GB
    Quality evidence
    • held-out test, 8 synthetic incidents: planted additions / contradictions / omissions found16/16 / 16/16 / 16/16decosa-api docs/evals/report-integrity.md, measured on our server 2026-09-25, gateway route, one run; data written by the building agent, prompts frozen on a separate 4-case dev split
    • faithful reports: sentences flagged / key events flagged as left out0/64 / 9/68 (2/68 as missing, 7 as partly)decosa-api docs/evals/report-integrity.md, measured on our server 2026-09-25, gateway route
    • the two demo incidents on real speech-recognition transcripts4/4 additions, 4/4 contradictions, 4/4 omissions found; faithful sentences flagged 2/19 (1 partial, 1 contradicted)decosa-api docs/evals/report-integrity.md (asr run), measured on our server 2026-09-26, gateway route: the traffic-stop and bar-fight reports checked against MOSS-Transcribe-Diarize transcripts of their synthetic audio, re-voiced 26 Sep 2026 from macOS voices to Decosa house voices (Kokoro-82M) (the first build: 1/19 flagged, partial)
    • real reports against real body-worn-camera audio, reviewed by a lawyernot measured yet
    Latency
    measured on our server: seconds to half a minute per check on a quiet gateway, up to a few minutes while the shared gateway was busy
    Verification
    Proof: strongHosted: gateway-signed receipt per model call. Self-hosted: attested by the box's key.
  • Best

    DeepSeek-V4-Flash on two more cards

    A larger judge for long, many-speaker incidents.

    Models
    • MOSS-Transcribe-Diarize 0.9B
    • DeepSeek-V4-Flash (NVIDIA NVFP4)
    Hardware
    2x RTX PRO 6000 96 GB
    Quality evidence
    • planted discrepancies found / faithful sentences flaggednot measured yet
    Latency
    not measured yet
    Verification
    Proof: partialSelf-host onlyNot a hosted model: calls are attested by the box's key only.
  • Needs more compute

    Wanted: the best setup

    two large judges from different families

    DeepSeek-V4-Flash and GLM-5.3-Flash each read the file, and a finding stands when they agree; disagreements go to the reviewer. Reports and transcripts stay on your own hardware, never on community providers. Not served yet.

    Models
    • MOSS-Transcribe-Diarize 0.9B
    • DeepSeek-V4-Flash (NVIDIA NVFP4)
    • GLM-5.3-Flash
    Hardware
    Your own hardware: 2x 96 GB cards for DeepSeek-V4-Flash plus 2x 96 GB for GLM-5.3-Flash, or one Mac Studio with 512 GB holding both 4-bit builds (156 + 165 GB, sizes from our Mac; not run together yet). Estimate.
    Quality evidence
    • planted discrepancies found / faithful sentences flaggednot measured yet
    Latency
    not measured yet
    Verification
    No proof yetSelf-host onlyOn your own hardware its calls are attested by the box's key only: not a hosted model there, so no gateway receipts. Never sent to community providers.
    Not served yet. It needs more than one 96 GB card, so it runs on your own bigger box.
Components

Every model in the stack

Models in this stack. Each row has a button that shows its licence, engine, verification and evidence.
ModelDetails
After the session
Key events, first-draft writer, sentence judge (the grounding checker) and event coverageQwen3.8-27B (NVIDIA NVFP4)nvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab)
27.8B · 57 GBProof: strongIn the hosted demo
Recording to a timed, speaker-labelled transcript (POST /report/transcribe, self-host)MOSS-Transcribe-Diarize 0.9BOpenMOSS-Team/MOSS-Transcribe-Diarize on Hugging Face (opens in a new tab)
0.9BProof: partialSelf-host only
Lite tier: the same four steps on a 48 GB cardGemma 4 26B A4B (instruction-tuned)google/gemma-4-26B-A4B-it on Hugging Face (opens in a new tab)
25.2B (3.8B active)No proof yetSelf-host only
Best tier: a larger judge for long, many-speaker incidentsDeepSeek-V4-Flash (NVIDIA NVFP4)nvidia/DeepSeek-V4-Flash-NVFP4 on Hugging Face (opens in a new tab)
284B (13B active) · 192 GBNo proof yetSelf-host only
Other
Second judge, from another familyGLM-5.3-Flashzai-org/GLM-5.3-Flash on Hugging Face (opens in a new tab)
321B (18B active) · about 170 GB (estimate)No proof yetSelf-host only

Around the models

Tools, services and hardware

Tools

Services

  • decosa-api:8445
    ${DECOSA_REGISTRY}/decosa-api:0.1.0

    Transcript intake, the check and draft runs, finalize, signing and the HTTP API (/report/*). No GPU. Binds 127.0.0.1 by default.

  • decosa-llm:8000
    ${DECOSA_REGISTRY}/decosa-llm:0.1.0

    vLLM OpenAI endpoint for Qwen3.8-27B. Internal to the compose network.

  • decosa-diarize:8092

    Optional, self-host only: MOSS-Transcribe-Diarize for transcripts made from recordings. No published image yet; built from services/diarize.

Hardware

  • 1x RTX PRO 6000 Blackwell 96 GB Fits

    Measured: the hosted demo's Qwen3.8-27B runs on one of these cards on our server; the diarizer runs on the other card there.

  • 1x L40S / RTX 6000 Ada 48 GB

    Not measured. FP8 Qwen3.8-27B with a shorter context, or Gemma 4 26B A4B (lite).

  • CPU only Fits

    Transcript parsing, finalize (diffs, provenance, the signed record) and verification need no GPU; the check itself needs the model.

Latency per lane

  • check with key events, self-hosted (direct route, own vLLM)6.0 s

    Measuredmeasured on our server 2026-09-25, self-host sandbox against the local Qwen3.8-27B vLLM; 13 s for a draft with its check

  • check of a 10-12 sentence report with key events, quiet gateway15.0 s

    Measureddecosa-api docs/evals/report-integrity.md, dev runs on our server 2026-09-25 (8-36 s), gateway route

  • check of an 8-11 sentence report with key events, busy shared gateway83.0 s

    Measureddecosa-api docs/evals/report-integrity.md, test runs on our server 2026-09-25 (43-168 s, mean 83 s), gateway shared with other evaluation jobs

  • transcribe a 100 s recording (self-host)5.9 s

    Measuredmeasured on our server 2026-09-25, MOSS-Transcribe-Diarize on an RTX PRO 6000

  • finalize (diff, provenance, signed record)50 ms

    Estimateestimate: no model call

Assemble it

Run this exact stack on your machine

Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.

report-integrity/assemble-prompt.md154 lines
# Assemble Decosa report integrity on this machine

You are setting up a self-hosted report-integrity checker on this Linux machine. It takes the transcript of a recording (body-worn camera, security camera, EMS or workplace interview) and a written incident report, checks every sentence of the report against the transcript, lists key events on the recording the report leaves out, and signs the result. It can also draft a report from the transcript, keep a signed first draft, and seal a record of which sentences a person changed before signing. With the optional diarizer it turns a recording into a timed, speaker-labelled transcript. Work step by step, show me each command before running anything that needs sudo, and stop if a check fails.

**Before anything else, remind me:** incident reports, body-worn-camera audio and discovery hold personal and often confidential criminal-justice information. Keep everything on this machine: the model route stays local (`direct`) and nothing goes to a hosted service. "Not in the recording" means the transcript does not contain it, not that it is false. This is not legal advice; disclosure and certification duties (California Penal Code 13663 from SB 524, in force 1 Jan 2026; Utah Code 53-25-602 from SB 180, in force 7 May 2025) are for the agency and its counsel. Repeat these points in your final summary.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/report-integrity.zip (4 KB, 13 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py report-integrity` (the api image carries the same bundle under /app/rehearsal/report-integrity/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py report-integrity --bundle report-integrity.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "every sentence got a verdict (none errored)", "'thirty feet' (the recording says ten) is flagged", "'the backup alarm was not working' (the recording says it was) is flagged"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## What you are building

| service | image | model | port |
|---|---|---|---|
| `llm` | `${DECOSA_REGISTRY}/decosa-llm:0.1.0` (vLLM 0.29.0, `vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1`) | `nvidia/Qwen3.8-27B-NVFP4` @ `482ca0f3832238542f8f5295dde86b5f22711d80`, Apache-2.0 | internal 8000 |
| `api` | `${DECOSA_REGISTRY}/decosa-api:0.1.0` (no GPU) | none | `127.0.0.1:8445` |
| `diarize` (optional, recordings only) | built from the decosa-api source (`services/diarize`) | `OpenMOSS-Team/MOSS-Transcribe-Diarize` @ `704aa4a9c304e8520be88901e0d1960158ef5b15`, Apache-2.0 | internal 8092 |

Typed, WebVTT or SRT transcripts need no speech model. `diarize` only turns a recording into a transcript.

## 1. Check the GPU, driver and Docker

1. Run `nvidia-smi`. I need one NVIDIA GPU with at least 48 GB and driver 580 or newer.
   - Blackwell (RTX PRO 6000, B200): use the defaults below (NVFP4). This is the measured setup.
   - Hopper (H100/H200) or 48 GB Ada/L40S: set `LLM_MODEL=Qwen/Qwen3.8-27B-FP8`, `LLM_REVISION=main`; on 48 GB also `LLM_MAX_LEN=32768`. Not measured.
   - Under 48 GB: stop and tell me it will not fit.
2. Check `docker --version`, `docker compose version` and `docker run --rm --gpus all ubuntu nvidia-smi`. If Docker or the NVIDIA Container Toolkit is missing, install them from the official Docker and NVIDIA repositories (`sudo nvidia-ctk runtime configure --runtime=docker`, then restart Docker).
3. Confirm about 60 GB of free disk.

## 2. Get the images

The images are **on request** while self-host is in early access: ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry (set it as `DECOSA_REGISTRY`), pull access and the compose file. Try `docker pull ${DECOSA_REGISTRY}/decosa-{llm,api}:0.1.0`. If a pull fails, build from source once the `decosa-api` source is published: clone it and run `docker build -f docker/api/Dockerfile -t ${DECOSA_REGISTRY}/decosa-api:0.1.0 .` in it (and `docker compose build llm` from its compose file). If neither works, stop and tell me.

## 3. Write the compose file

Create `~/decosa-report/.env`:

```bash
DECOSA_TAG=0.1.0
DECOSA_GPU=0
LLM_MODEL=nvidia/Qwen3.8-27B-NVFP4
LLM_REVISION=482ca0f3832238542f8f5295dde86b5f22711d80
LLM_MAX_LEN=65536
LLM_GPU_UTIL=0.85
DECOSA_SIGNER_NAME="<who signs these checks, e.g. County Public Defender, investigations>"
```

Create `~/decosa-report/docker-compose.yml` with exactly these services:

```yaml
name: decosa-report
x-health: &health
  interval: 15s
  timeout: 5s
  retries: 5
services:
  llm:
    image: ${DECOSA_REGISTRY}/decosa-llm:${DECOSA_TAG}
    deploy: { resources: { reservations: { devices: [ { driver: nvidia, device_ids: ["${DECOSA_GPU:-0}"], capabilities: [gpu] } ] } } }
    ipc: host
    restart: unless-stopped
    volumes: [hf-cache:/root/.cache/huggingface]
    command: ["${LLM_MODEL}", "--revision", "${LLM_REVISION}", "--served-model-name", "qwen3.8-27b",
              "--language-model-only", "--max-model-len", "${LLM_MAX_LEN}", "--gpu-memory-utilization", "${LLM_GPU_UTIL}",
              "--max-num-seqs", "16", "--kv-cache-dtype", "fp8_e4m3", "--speculative-config", '{"method":"mtp","num_speculative_tokens":3}',
              "--seed", "0", "--enable-force-include-usage", "--disable-uvicorn-access-log", "--host", "0.0.0.0", "--port", "8000"]
    healthcheck: { <<: *health, test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8000/health', timeout=4)"], start_period: 900s }
  api:
    image: ${DECOSA_REGISTRY}/decosa-api:${DECOSA_TAG}
    restart: unless-stopped
    depends_on: { llm: { condition: service_healthy } }
    environment:
      DECOSA_LLM_ROUTE: direct                  # local model only; receipts are signed by this box's key ("attested")
      DECOSA_LLM_URL: http://llm:8000/v1
      DECOSA_LLM_MODEL: qwen3.8-27b
      DECOSA_LOCAL_SIGNING: "on"                # Ed25519 key created at /data/attest/ed25519.pem on first start
      DECOSA_SIGNER_NAME: ${DECOSA_SIGNER_NAME}
      DECOSA_SESSIONS_PER_IP_HOUR: "1000"
      DECOSA_BUDGET_LLM_TOKENS: "200000"        # per session; a 15-sentence check with key events needs about 5,000
      DECOSA_SESSION_TTL_S: "28800"
      DECOSA_REPORT_AUDIO: "0"                  # "1" only with the diarize service
      DECOSA_DIARIZE_URL: ""
      DECOSA_CORS_ORIGIN_REGEX: '^https?://(localhost|127\.0\.0\.1)(:\d+)?$$'
    ports: ["127.0.0.1:8445:8445"]
    volumes: [decosa-data:/data]
    healthcheck: { <<: *health, test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8445/healthz', timeout=4)"], start_period: 20s }
volumes: { hf-cache: {}, decosa-data: {} }
```

Use the named volume `decosa-data` as written (a host bind mount owned by root makes the API fail on `/data/keys.sqlite`). Run `docker compose up -d` and poll `docker compose ps` until both are healthy (the LLM takes 5-10 minutes the first time). `curl -s localhost:8445/healthz` should show `"llm": true` (`asr` is false here and that is fine).

## 4. Optional: recordings (diarize)

Only if I want transcripts made from recordings: build `services/diarize` from the decosa-api source on a CUDA PyTorch base image (env `DIARIZE_HOST=0.0.0.0 DIARIZE_PORT=8092 DIARIZE_DEVICE=cuda:0`, the `hf-cache` volume, the same GPU, a health check on `GET /health`), lower `LLM_GPU_UTIL` to 0.80, and set `DECOSA_DIARIZE_URL: http://diarize:8092` and `DECOSA_REPORT_AUDIO: "1"`. `POST /report/transcribe` then takes a 16 kHz mono 16-bit WAV (`ffmpeg -i in.mp4 -ac 1 -ar 16000 -sample_fmt s16 out.wav`) and returns timed segments with a speech receipt; pass them as `segments` (with `speakers` to name S01, S02) to `/report/check` or `/report/draft`. The fit beside the LLM is an estimate, not measured.

## 5. Smoke test

```bash
API=localhost:8445
TOKEN=$(curl -s $API/demo/session -H 'content-type: application/json' -d '{"vertical":"report-integrity"}' | jq -r .token)
curl -s $API/report/check -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' \
  -d '{"sample_id":"traffic-stop-audio","report":{"sample":"tampered"}}' > /tmp/check.json
jq '{counts, omission_counts, missing, receipts: (.receipts|length)}' /tmp/check.json
jq -r '.sentences[] | "\(.verdict)\t\(.text)"' /tmp/check.json
jq '{statement: .signed}' /tmp/check.json | curl -s $API/report/verify -H 'content-type: application/json' -d @- | jq .
```

Pass if: the 62 mph and "did not have proof of insurance" sentences are `contradicted`, the odor and center-console sentences are `unsupported`, key events about the one beer and the refused search are `missing`, every receipt has `"status": "attested"`, and the verify shows `valid_signature` and `signed_by_this_server` true. Then the faithful report, `{"sample":"faithful"}`, should come back with few or no flags. On the measured card a check takes 30-80 s depending on load.

Then the draft and record path:

```bash
curl -s $API/report/draft -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' -d '{"sample_id":"bar-fight-audio"}' > /tmp/draft.json
jq '{first: .first_draft.id, sentences: (.draft.sentences|length), to_confirm: .draft.to_confirm}' /tmp/draft.json
jq '{first_draft, draft: .draft.text, versions: [{text: (.draft.text + " I did not see who threw the first punch.")}], jurisdiction: "CA", certification: {name: "Test Author", accepted: true}}' /tmp/draft.json \
  | curl -s $API/report/finalize -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' -d @- > /tmp/final.json
jq '{counts, disclosure: .disclosure.text, self_check}' /tmp/final.json
jq '{record}' /tmp/final.json | curl -s $API/record/verify -H 'content-type: application/json' -d @- | jq '{ok, summary}'
```

Pass if the counts show one `human` sentence, the disclosure starts "This report was written either fully or in part using artificial intelligence.", and the record verifies (`ok: true`).

## 6. Point the app at the local API

- Base URL `http://localhost:8445` (web app: `NEXT_PUBLIC_DECOSA_API=http://localhost:8445`). Add other origins to `DECOSA_CORS_ORIGINS`.
- `POST /report/check` and `/report/draft` return JSON, or stream Server-Sent Events with `Accept: text/event-stream`. Keep the signed statement (and in draft mode the signed first draft) with the report; they hold hashes, not text.
- Keep the API on 127.0.0.1. For other users on the LAN, put a TLS reverse proxy with authentication in front and set `DECOSA_TRUSTED_PROXIES`.

## 7. Keep the direct route

`DECOSA_LLM_ROUTE=gateway` would send prompts, which contain the transcript and the report, to the hosted Decosa
API. Never use it for real incidents. Synthetic training material at most.

Finish with a summary: what is running, the health output, the smoke-test results, and the reminders above.
Rules and regulations it checks againstDated, linked to the primary source; not legal advice

Regulation watch

Loading the watch status…

1 law, rule and guidance page cited; 1 watched nightly at the primary source. A change marks this page for a human re-check; nothing is edited automatically. What we cite and how it is watched

Technical detailsModels, where it runs, labels

In short

Last reviewed

What it is
An AI police report check: every sentence of an incident report, written by an officer or drafted by AI, is compared with the transcript of the body cam or other recording it describes, with the times, and the key events the report leaves out are listed. It runs on your own GPU.
Who it's for
Teams in public sector and legal.
Where it runs
Self-host (hosted demo: synthetic material only)
Key numbers

On 8 held-out synthetic incidents (test split, one run) it found 16 of 16 planted additions, 16 of 16 contradictions and 16 of 16 omissions, and flagged 0 of 64 faithful sentences. The building agent wrote the scenarios, so they are clear-cut; it has not been measured on real body-worn-camera audio.

  • 16/16 Planted additions found (test split, n = 16)
  • 16/16 (16/16) Planted contradictions found (exact) (test split, n = 16)
  • 16/16 (13/16) Planted omissions found, missing or partly (missing only) (test split, n = 16)
  • 97.0 s Median end-to-end run, hosted (QA sweep 2026-09-25)
All results, datasets and caveats
Models
MOSS-Transcribe-Diarize · Qwen3.8-27B
Where
Self-host (hosted demo: synthetic material only)
Checks
Receipt per model call; signed check; signed first draft and hash-chained edit record
Industry
Public sector · Legal
Output
Signed record or verdict · Notes, reports and drafts
Data
Personal data · Confidential business data
Hardware
1× 96 GB GPU
Licence
Permissive (Apache-2.0, MIT)
Runs in
Self-host

Questions people ask

What does the AI police report check return?

Each report sentence comes back as in the recording, partly, not in the recording, or the recording says otherwise, with the transcript lines and times it rests on. Key events on the recording (force, injuries, rights, consent, custody, what each person said) are looked for in the report, and missing ones are listed with times you can cue up.

Can AI write a police report?

Agencies already draft reports with AI, and two states regulate it. Draft mode here writes a first draft from the recording, signs it, and seals a record of which sentences the officer changed before signing, with the disclosure line. The officer still verifies and signs the facts. The quality of the first draft is not measured; the check itself is what the numbers on this page describe.

Does 'not in the recording' mean the report is false?

No. It means the transcript does not contain it. A camera can miss what a person saw or smelled, so those sentences are flagged for the author's own account. The check also trusts the transcript: a mishearing turned into a fact is marked as in the recording.

How well does it work?

On 8 held-out synthetic incidents (police, security, EMS and workplace) it found 16 of 16 planted additions, 16 of 16 contradictions and 16 of 16 omissions (13 of 16 as fully missing), and flagged 0 of 64 faithful sentences. The scenarios were written by the building agent and are clear-cut; it has not been measured on real body-worn-camera audio.

Does it help with California SB 524 or Utah SB 180?

As a record-keeping aid. California SB 524 (in force 1 Jan 2026) requires a report written partly with AI to say so and name the program, carry the officer's signature verifying the facts, keep the first draft as long as the report, and keep an audit trail of who used AI and the footage used. Utah SB 180 requires a disclaimer and the author's certification. Draft mode signs the first draft and seals which sentences a person changed, with the disclosure line; the text follows the statutes but is not legal advice.

Can we run it on real discovery?

Run it on your own GPU: real reports and footage are personal and often confidential, and the stack has no CJIS attestation. Audio intake is self-host only, and the hosted demo takes synthetic material only.

Ask a question or leave feedbackWe read every message and publish useful answers
Questions & feedback

Ask about Report integrity

We read every message. Questions, comments and our answers show here once we have reviewed and approved them.

Loading questions…

This is a

Plain text. Please leave out personal, patient or client data.

Shown with your message if we publish it. Leave blank to post as “A visitor”.

Nothing appears here until we have read and approved it.