Write a SAR narrative
A SAR narrative in FinCEN's structure where every sentence cites its rows and every amount, date and count is recomputed from them.
Built on: Numeric grounding, Grounding, Signed record
Loading the tool…
Use it your way
Use it from your codeThe hosted API with your key, and prompts to paste into a coding agent
Run it yourself, on request
- The same open models and app, on 1× RTX PRO 6000 (96 GB) or 1× RTX 5090 (32 GB) for Qwen3.8-27B; the numeric check, the screen and the report run on CPU.
- Data never leaves your machines, and there are no Decosa charges.
- One prompt for Claude Code or Codex assembles the whole stack.
- Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Get an API key
- Call the sar narrative desk API from your own code in minutes.
- Every model answer carries a signed receipt.
- Synthetic, public or test data only: real confidential data belongs on your own hardware.
Build with it
Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.
- Base URL
- https://api.decosa.ai
- Auth
Authorization: Bearer $DECOSA_API_KEY(or a demo session token)- Tool id
- sar-narrative-desk
Use the hosted API
# Decosa SAR narrative desk: use the hosted API (synthetic case files only)
You are wiring Decosa's SAR narrative desk into this project. From an alert's case file (a transactions CSV, customer
profile and alert fields, the investigator's notes) it drafts a suspicious-activity narrative in FinCEN's structure (who,
what, when, where, why, how), or a no-file rationale, or checks a narrative the investigator wrote. Every sentence cites
the rows, fields or notes it rests on; every amount, date, count and total is recomputed in code from the cited rows;
the other facts are judged against only the cited evidence. It returns a filing copy, a cited workpaper and a report
signed by the server, with a signed receipt for every model call. Use only what is listed below. If you need something
else, stop and ask me.
- Base URL: `https://api.decosa.ai`
- Health check: `GET https://api.decosa.ai/healthz`.
- **The hosted API takes synthetic case files only.** SAR information is confidential (31 U.S.C. 5318(g)(2); 31 CFR
1020.320(e)). Every request must carry `"synthetic": true`; anything else gets HTTP 400. Real cases belong on a
self-hosted box (see the self-host prompt). Never send real names, accounts or transactions here.
- This is a drafting aid, not a filing and never a compliance determination. Say so wherever you show results; an
investigator decides whether to file and files.
## Auth: API key (or a demo session)
1. Preferred: an API key (`dk_…`) from "Get an API key" on the tool page. Keep it in an environment variable,
`DECOSA_API_KEY`, never in code. Send `Authorization: Bearer $DECOSA_API_KEY`.
2. Without a key: `POST https://api.decosa.ai/demo/session` with `{"vertical": "sar-narrative-desk"}` returns `{"token", "expires_at", "budget"}`.
Sessions per IP are limited; over a limit you get HTTP 429 with `Retry-After`.
3. A full draft needs about 7,000 generated tokens of budget (402 otherwise); a check needs about 260 per sentence. One
run at a time per demo token (409 while one is going).
## Endpoints
- `POST /sar/draft` (token). Body:
```json
{"transactions": "id,date,type,direction,amount,channel,location,counterparty,memo\nT1,2026-03-04,cash deposit,in,9400.00,branch,Main St branch,,",
"kyc": {"name": "...", "occupation": "...", "annual_income": "$38,000.00", "expected_activity": "..."},
"alert": {"case_id": "...", "alert_rule": "...", "alert_date": "2026-03-20", "institution": "..."},
"notes": "The investigator's notes.", "mode": "draft | no_file | check",
"narrative": "only for mode check: sentences ending with [T3, K.occupation, N2]", "synthetic": true, "stream": false}
```
CSV needs `date` and `amount` columns (other columns optional); ≤ 400 rows, ≤ 120,000 characters. Rows are cited
`T1..Tn` (or the CSV's own `T` ids); profile fields `K.<name>`; alert fields `A.<name>`; note sentences `N1..`; the
screen's computed facts `F1..`.
- JSON answer: `{status: held | needs_review | all_checked, counts, numbers: {ok, mismatch, unverified}, typologies,
warnings, screen: {facts, typologies, indicated, possible}, sections, sentences: [{sid, section, text, clean, cites,
status: checked | flagged | held | not_checked | no_claim, reasons, numeric: {mentions: [{text, kind, status,
derivation | detail}]}, grounding, receipt_ids}], narrative_md (workpaper), narrative_plain (filing copy), report,
receipts, budget}`. Show `held` sentences with their `reasons` and the figure in `detail`; never show a held sentence
as part of the filing copy.
- SSE: send `Accept: text/event-stream` (or `"stream": true`): `ready`, `section`, `receipt`, `sentence`, `report`,
`budget`, `done`.
- `POST /sar/numbers` (token): the same body with `narrative`; the numeric check only, no model call.
- `POST /sar/verify` (no token): `{report, narrative_md?, narrative_plain?, transactions?}` → `{valid_signature,
signed_by_this_server, ...matches}`.
- `GET /sar/info`, `GET /sar/samples` (no token): the checks, thresholds, limits, sources, and six synthetic samples.
## Errors
400 bad input (the message names the field or the limit, or says the case is not marked synthetic), 401/403 token,
402 budget, 409 a run already going on this demo token, 413 body over 512 KB, 429 busy (`Retry-After`).
Run it yourself (containers)
On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.
# Decosa SAR narrative desk: run it yourself (containers)
You are setting up the Decosa SAR narrative desk on this machine, so case files never leave it. SAR information is
confidential (31 U.S.C. 5318(g)(2); 31 CFR 1020.320(e)). From a transactions CSV, the customer profile and the
investigator's notes it drafts a cited suspicious-activity narrative (or a no-file rationale, or checks a draft),
recomputes every number from the cited rows, and returns a filing copy, a cited workpaper and a signed report. Nothing
is sent to Decosa's hosted API. It is a drafting aid: an investigator decides and files.
Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.
Ask me before any command that needs sudo, and show me the command first.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/sar-narrative-desk.zip (4 KB, 15 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py sar-narrative-desk` (the api image carries the same bundle under /app/rehearsal/sar-narrative-desk/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py sar-narrative-desk --bundle sar-narrative-desk.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the run holds sentences back", "the wrong total is held as a number mismatch", "the held total names the figure the rows give"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
instructions for this distribution (docs.docker.com/engine/install). Install the NVIDIA container toolkit and check
`docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
`mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
Read it. Keep the `llm` service (Qwen3.8-27B on vLLM, with prefix caching on) and the `api` service. For the `api`
service set `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b`,
`DECOSA_SAR_SYNTHETIC_ONLY=0` (so this box accepts real cases) and bind every port to 127.0.0.1. Never set the
gateway route on this box: it would send case data to the Decosa API.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check (the first start
downloads about 20 GB of weights).
4. Check: `curl -fsS http://127.0.0.1:<PORT>/sar/info` lists the modes, the checks, the screen thresholds, the limits
and `synthetic_only: false`; `GET /attest/signing-key` shows this box's public key. Show me the key: it is what a
reviewer pins to verify my reports.
5. Smoke test: get a token with `POST /demo/session {"vertical":"sar-narrative-desk"}`, fetch `GET /sar/samples`, and
send the `planted-errors` sample (transactions, kyc, alert, notes, mode, narrative, synthetic) to `POST /sar/draft`.
Expect `status: "held"`: the wrong total and the shifted date held as `number_mismatch` with the right figures, and
the invented $25,000.00 wire held. Then send the `structuring` sample (mode `draft`): expect eight sections and
`screen.indicated: ["structuring"]`. Then `POST /sar/verify` with the report, narrative_md and narrative_plain:
`valid_signature`, `signed_by_this_server` and both matches must be true.
6. Report back: the public key, both statuses, the held sentences and how long each run took.
Off by default. Joining as a provider serves other people's requests on this GPU; never do it on a box that holds SAR
case data. If I ask for it later, follow the Provide page instead of improvising.
No NVIDIA GPU? This tool also runs entirely on an Apple Silicon Mac (MLX, 32 GB of unified memory or more): use https://decosa.ai/prompts/sar-narrative-desk-mac.md instead.
Run it on your own hardwareWhat it needs, and the prompt that sets it up
Run it on your own GPU
Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.
Hardware check
Check your own hardware- CPU only, 64 GB RAMlite tierRuns with a smaller tier
The standard tier does not fit: Qwen3.8-27B (NVFP4) needs a GPU. The lite tier fits.
- GeForce RTX 4090standard tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)
- GeForce RTX 5090standard tierRuns
The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
- 2x GeForce RTX 5090standard tierRuns
The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size).
- L40Sstandard tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.
- H100 80 GB (SXM)standard tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.
- RTX PRO 6000 Blackwell 96 GBstandard tierRuns
The standard tier fits (57.6 of 96 GB).
- 2x RTX PRO 6000 Blackwell 96 GBstandard tierRuns
The standard tier fits (57.6 of 192 GB).
- Apple M3 Ultra (Mac Studio), 96 GBstandard tierRuns
The standard tier fits (32 of 96 GB).
- Apple M5 Max, 64 GBstandard tierRuns
The standard tier fits (32 of 64 GB).
Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.
On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.
- 1
Check the GPU, Docker and the NVIDIA Container Toolkit
The driver must see the GPU, and Docker must be able to pass it into a container.
nvidia-smi docker compose version docker run --rm --gpus all ubuntu nvidia-smi
- 2
Fetch the compose file
One file describes the API and the language model as services.
mkdir -p ~/decosa && cd ~/decosa curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml - 3
Pull and start
The first start downloads pinned model weights, tens of gigabytes.
docker compose pull docker compose up -d
- 4
Check health
Wait until the API reports ok with the language model loaded. Then point your app at the local base URL.
curl -fsS http://localhost:<PORT>/healthz # {"ok": true, "llm": true, ...} curl -fsS -X POST http://localhost:<PORT>/demo/session \ -H 'Content-Type: application/json' -d '{"vertical":"sar-narrative-desk"}'
Set up with a coding agent, rehearse on mock data, then go private
- Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
- Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's
expected.json. Every check must print PASS. - Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
docker compose exec api python scripts/rehearse.py sar-narrative-desk
Download the mock-data bundle (4 KB, 15 checks)expected.json
A synthetic structuring case (a transactions CSV, KYC and alert fields, investigator notes) and an investigator draft with citations, in which three things were planted: a total $1,000 too high, a date moved by two days, and a $25,000 wire that is not in the ledger. Check mode makes no drafting call: every number is recomputed from the cited rows and each sentence is judged against what it cites. The three planted sentences must be held, the right figures named, the screen must find structuring, the filing copy must drop the held sentences and the citations, and the signed report must verify and catch a changed status.
What the rehearsal checks
- the run holds sentences back
- the wrong total is held as a number mismatch
- the held total names the figure the rows give
- the shifted date is held as a number mismatch
- the invented wire is held
- the invented wire is named as not in the case file, not as a number mismatch
- two numbers do not match and one is not in the case file at all
- the invented amount is counted as not found
- the screen finds structuring
- the filing copy drops the invented wire
- the filing copy has no citation brackets
- the signed report verifies
- the report matches the ledger it was run on
- a report with its status changed to all_checked no longer verifies
- every model call has a signed receipt
Licence: Synthetic: every person, account, business and the bank (Example Community Bank) are fictional, generated by decosa_api/verticals/sar/synth.py. Part of decosa-api, AGPL-3.0-or-later.
Prompt for your coding agent
# Decosa SAR narrative desk: run it yourself (containers)
You are setting up the Decosa SAR narrative desk on this machine, so case files never leave it. SAR information is
confidential (31 U.S.C. 5318(g)(2); 31 CFR 1020.320(e)). From a transactions CSV, the customer profile and the
investigator's notes it drafts a cited suspicious-activity narrative (or a no-file rationale, or checks a draft),
recomputes every number from the cited rows, and returns a filing copy, a cited workpaper and a signed report. Nothing
is sent to Decosa's hosted API. It is a drafting aid: an investigator decides and files.
Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.
Ask me before any command that needs sudo, and show me the command first.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/sar-narrative-desk.zip (4 KB, 15 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py sar-narrative-desk` (the api image carries the same bundle under /app/rehearsal/sar-narrative-desk/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py sar-narrative-desk --bundle sar-narrative-desk.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the run holds sentences back", "the wrong total is held as a number mismatch", "the held total names the figure the rows give"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
instructions for this distribution (docs.docker.com/engine/install). Install the NVIDIA container toolkit and check
`docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
`mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
Read it. Keep the `llm` service (Qwen3.8-27B on vLLM, with prefix caching on) and the `api` service. For the `api`
service set `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b`,
`DECOSA_SAR_SYNTHETIC_ONLY=0` (so this box accepts real cases) and bind every port to 127.0.0.1. Never set the
gateway route on this box: it would send case data to the Decosa API.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check (the first start
downloads about 20 GB of weights).
4. Check: `curl -fsS http://127.0.0.1:<PORT>/sar/info` lists the modes, the checks, the screen thresholds, the limits
and `synthetic_only: false`; `GET /attest/signing-key` shows this box's public key. Show me the key: it is what a
reviewer pins to verify my reports.
5. Smoke test: get a token with `POST /demo/session {"vertical":"sar-narrative-desk"}`, fetch `GET /sar/samples`, and
send the `planted-errors` sample (transactions, kyc, alert, notes, mode, narrative, synthetic) to `POST /sar/draft`.
Expect `status: "held"`: the wrong total and the shifted date held as `number_mismatch` with the right figures, and
the invented $25,000.00 wire held. Then send the `structuring` sample (mode `draft`): expect eight sections and
`screen.indicated: ["structuring"]`. Then `POST /sar/verify` with the report, narrative_md and narrative_plain:
`valid_signature`, `signed_by_this_server` and both matches must be true.
6. Report back: the public key, both statuses, the held sentences and how long each run took.
Off by default. Joining as a provider serves other people's requests on this GPU; never do it on a box that holds SAR
case data. If I ask for it later, follow the Provide page instead of improvising.
No NVIDIA GPU? This tool also runs entirely on an Apple Silicon Mac (MLX, 32 GB of unified memory or more): use https://decosa.ai/prompts/sar-narrative-desk-mac.md instead.
Help me customise for my hardware
Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.
GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page
RunsSAR narrative desk on GeForce RTX 5090: use the Standard · one GPU for the model (hosted demo) tier
The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
What this tool's stack says about this hardware:
- 1x RTX 5090 32 GB (fits): Estimate: Qwen3.8-27B NVFP4 needs about 20 GB of weights plus KV cache; not run for this use case.
Standard · one GPU for the model (hosted demo): what changesuses estimates
- Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
Memory per component
- Red-flag screen, citation check, numeric grou...: decosa-api SAR desk (decosa_api/verticals/sar) with the numeric grounding block (decosa_api/verticals/numeric) and the grounding module (decosa_api/verticals/grounding). CPU. Runs on CPU (vram_gb 0 in stack.json).
- Section drafting and the grounding judge: Qwen3.8-27B (NVFP4). ~57.6 GB (at least ~28 GB), weights 21.4 GB (from stack.json). Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)
Expected speed
Not measured.
Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.
Setup prompt for this hardware
The self-host prompt for SAR narrative desk, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.
# Set up SAR narrative desk on my hardware Fetch https://decosa.ai/prompts/sar-narrative-desk-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied. ## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=sar-narrative-desk) Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4). Quality tier: Standard · one GPU for the model (hosted demo) (standard). Fit check: runs with changes, about 28 GB of 32 GB used; some memory numbers are estimates, not measurements. First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything. Use these components (the setup below describes the standard tier; change it to match): - Red-flag screen, citation check, numeric grou...: decosa-api SAR desk (decosa_api/verticals/sar) with the numeric grounding block (decosa_api/verticals/numeric) and the grounding module (decosa_api/verticals/grounding), CPU - Section drafting and the grounding judge: Qwen3.8-27B (NVFP4) (nvidia/Qwen3.8-27B-NVFP4), 57.6 GB. Change: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions. GPU placement (set each service's device and its vLLM --gpu-memory-utilization to about the share shown): - GPU 0: Qwen3.8-27B (NVFP4) ~28 GB (88%); about 4 GB left During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed. The stack's own component list and compose layout: https://decosa.ai/prompts/sar-narrative-desk-assemble.md
Or on a Mac Studio
No NVIDIA GPU needed: every model this tool uses runs natively on Apple Silicon through MLX. Any M-series Mac with 32 GB of unified memory or more. Measured speeds and what runs where
From a checkout of decosa-api, one command sets up the models and the API: scripts/mac/setup.sh
Mac prompt for your coding agent
# Decosa SAR narrative desk: run it on this Mac (Apple Silicon, no NVIDIA GPU) You are setting up the Decosa SAR narrative desk on this Mac, natively on Apple Silicon. The models run on the Mac's GPU through MLX and decosa-api runs from a git checkout with `uv`. Docker is not used for the models, because Docker on macOS cannot reach the GPU. Nothing is sent to Decosa's hosted API. Every model this tool needs runs on the Mac. It needs 32 GB of unified memory or more. Ask me before any command that needs sudo or installs software with Homebrew, and show me the command first. Never stop or kill a process this setup did not start; if a port is taken, pick another one. ## Step 0: set up with a coding agent, rehearse on mock data, then go private This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works. Work in this order: 1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to "test with something realistic". 2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool, https://decosa.ai/samples/sar-narrative-desk.zip (4 KB, 15 checks, synthetic or openly licensed: see `licence` in expected.json), show me what is in it, and run the rehearsal against the local API: `.venv/bin/python scripts/rehearse.py sar-narrative-desk` in the decosa-api checkout (the key comes from ~/.decosa-mac/api.key). It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the run holds sentences back", "the wrong total is held as a number mismatch", "the held total names the figure the rows give"). Show me the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json` to make a check pass. 3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this machine. For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent can read. Switch to your own data only after the rehearsal has passed and the agent's work is done. ## What runs where | Part | On an NVIDIA GPU | On this Mac | Status | |---|---|---|---| | Red-flag screen, citation check, numeric grounding, filing copy, workpaper and signed report (no model; CPU) | Python on CPU | The same Python module, run with uv | Runs, measured | | Section drafting and the grounding judge | NVFP4 on vLLM 0.29 (Blackwell) | MLX 4-bit (EigenLabs/Qwen3.8-27B-4bit) on mlx_lm.server 0.31.3; oMLX 0.6.1 with MTP as an option | Runs, measured | ## Steps 1. Check the machine: `uname -m` must print `arm64` (an M-series chip; Intel Macs cannot run MLX), and `sysctl -n hw.memsize` should be at least 32 GB for this tool. Check about 30 GB of free disk with `df -h ~`. Show me the chip (`sysctl -n machdep.cpu.brand_string`) and the memory. 2. Tools: `uv --version`. If it is missing, ask me, then `brew install uv`. 3. Code: `git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host> ~/decosa-api` (access required) and `cd ~/decosa-api`. Check that `scripts/mac/setup.sh` exists; if it does not, the checkout is too old: stop and tell me. 4. Start everything with one command: `scripts/mac/setup.sh`. It creates `.venv` (decosa-api) and `.venv-mac` (MLX, mlx-lm, mlx-audio), downloads the weights with the Hugging Face CLI (about 16 GB for the language model), starts the model servers and decosa-api on 127.0.0.1, and mints a local API key into `~/.decosa-mac/api.key` (mode 0600). The first run takes a while because of the downloads; later runs reuse them. If a download fails with 401 or 403, ask me for a Hugging Face token and set `HF_TOKEN`. 5. Check health: `scripts/mac/setup.sh status` shows each server, and `curl -fsS http://127.0.0.1:8445/healthz` must report `"llm": true`. `curl -fsS http://127.0.0.1:8445/attest/signing-key` shows this Mac's public key: show it to me, because it is what others pin to check the receipts and records this Mac signs. 6. Smoke test: `.venv/bin/python scripts/mac/bench_usecases.py sar-narrative-desk`. It runs the tool's own sample end to end against the local API with the local key and prints `ok`, the wall time, the model calls and the receipts. `ok=True` is the pass condition. If it fails, read `~/.decosa-mac/logs/*.log` and tell me what you found. 7. Point the app at it: the API is `http://127.0.0.1:8445` with `Authorization: Bearer $(cat ~/.decosa-mac/api.key)`, the same routes as the hosted API. To stop everything: `scripts/mac/setup.sh stop`. 8. Report back: the chip and memory, the public key, the smoke-test result and its time, and the output of `scripts/mac/setup.sh status`. ## Good to know - Receipts: every model call is signed with this Mac's own Ed25519 key and names the exact MLX weights (`qwen3.8-27b-mlx-4bit` with a hash of the downloaded files). There is no gateway countersignature on a self-hosted Mac. - The weights are a 4-bit MLX build of the same open models, not the NVFP4 build the hosted route and the published evals use. Expect small differences in wording and scores. - Faster drafting: `scripts/mac/setup.sh stop && scripts/mac/setup.sh --engine omlx` serves the model with oMLX and multi-token prediction (about 2x faster for a single long answer, no faster for many parallel calls; typed judgments then use sampling because oMLX returns no log-probabilities). - Built from the same parts as the measured tools; not run on the Mac yet. - Measured speeds for a Mac Studio M3 Ultra and the memory each tool needs: https://decosa.ai/mac. Full details: `docs/self-host-mac.md` in the checkout.
The proof
How we tested itEval results and end-to-end checks, hosted and self-hosted, with dates
Verified end to end
Hosted: verified 26 Sep 2026 · measured 26 Sep 2026: · p50 25 s · ~$0.017 per run · 30 receipts
Loading the nightly status…
Self-host: verified 26 Sep 2026 · fresh clone, compose up, sample against local model servers
Measured cost to run: about $0.017 per draft (hosted, 26 Sep 2026). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.
A fresh clone of a decosa-api pre-release build (not yet merged to main), the api image built from it with DECOSA_SAR_SYNTHETIC_ONLY=0, run against the already-running local Qwen3.8-27B vLLM on the direct route. The planted-errors check held exactly the three planted sentences (1.7 s); the structuring draft ran in 7.9 s with 51 of 51 numbers traced; the report verified, a changed status failed, and receipts were attested. Model-server startup itself not re-verified.
Known limits (5)
- Synthetic only on the hosted demo. Every number here comes from our own synthetic generator; it has not been run on real case files.
- A date is checked for membership in the cited rows, and tied to its amount when the sentence pairs them; a date moved onto another cited day with no amount beside it can pass. Counts can match another subset of the cited rows.
- The screen's thresholds are ours. It flagged structuring on 5 of 40 clean cash-business cases.
- It checks that what is written is traced to the case file, not that nothing is missing, and it sees nothing outside the case file (watch lists, other institutions, prior SARs).
- The ledger must be CSV; PDF statements are not read.
How it's builtThe steps, the models and what each one checks
Run it yourself, on request
- The same open models and app, on 1× RTX PRO 6000 (96 GB) or 1× RTX 5090 (32 GB) for Qwen3.8-27B; the numeric check, the screen and the report run on CPU.
- Data never leaves your machines, and there are no Decosa charges.
- One prompt for Claude Code or Codex assembles the whole stack.
- Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Get an API key
- Call the sar narrative desk API from your own code in minutes.
- Every model answer carries a signed receipt.
- Synthetic, public or test data only: real confidential data belongs on your own hardware.
A suspicious-activity narrative drafted from the case file, with every sentence citing its rows and every number recomputed in code.
For BSA/AML investigators. Send an alert's case file: the transactions as CSV, the customer profile and alert fields, and your notes. A red-flag screen in code (structuring, rapid movement, funnel account, elder exploitation) computes the totals, counts and date ranges with the rows behind them. Qwen3.8-27B drafts each section of the narrative in FinCEN's structure (introduction; who, what, when, where, why, how; conclusion), citing a row, field, note or computed fact for every sentence. Then each sentence is checked: its citations must exist; every amount, date, count, span and percentage is recomputed in code from the cited rows by the numeric grounding block; and the grounding judge reads it against only what it cites. A sentence that fails is held back from the filing copy, with the figure the rows do support. It also drafts a short no-file rationale for an alert closed without a SAR, or checks a narrative you wrote. You get a filing copy without citations, a cited workpaper for the case file, and a signed report. A drafting aid: an investigator decides and files.
- Deployment
- Self-host first
- Regulatory
- Checked 26 Sep 2026 against primary sources (links under Tools). SAR information is confidential: 31 U.S.C. 5318(g)(2) bars notifying anyone involved that a transaction was reported, and 31 CFR 1020.320(e) makes a SAR, and anything that would reveal one, confidential. So this is self-host first, and the hosted demo refuses any case file not marked synthetic. The draft follows FinCEN's Guidance on Preparing a Complete & Sufficient SAR Narrative (November 2003: the five W's plus how; introduction, body, conclusion; chronological; individual dates and amounts, not only totals; no tables or 'see attached') and the FinCEN SAR Filing Instructions (version 1.0, April 2026, Part V: do not file supporting documentation with the SAR), which is why the filing copy drops the citations and the cited workpaper stays in the case file (31 CFR 1020.320(d): keep supporting documentation five years). A bank files within 30 calendar days of initial detection, up to 60 if no suspect is identified (31 CFR 1020.320(b)(3)). The no-file rationale is optional: FinCEN's SAR FAQs of 9 October 2025 (Q4) say there is no requirement to document a decision not to file, FinCEN encourages it, and a short statement will usually suffice. Typologies follow FinCEN Advisories FIN-2014-A005 (funnel accounts, 28 May 2014) and FIN-2022-A002 (elder financial exploitation, 15 June 2022: an older adult is 60 or over); structuring follows 31 CFR 1010.100(xx) and FAQ Q1 (amounts near $10,000 alone do not require a SAR). The rapid-movement red flag cites the FFIEC BSA/AML manual appendix, which we could not re-fetch (unverified). The screen's thresholds are our own, not regulatory. This is not legal advice and never a compliance determination: the investigator decides whether to file, edits and files. Model licence: Apache-2.0 (Qwen3.8-27B).
Text description
An alert's case file (a transactions CSV, the customer profile and alert fields, and the investigator's notes, or a cited draft to check) goes to the SAR desk. A red-flag screen in code computes facts with their rows. Qwen3.8-27B drafts each section with citations, one receipted call per section. Each sentence is checked: its citations must exist, the numeric grounding block recomputes every number from the cited rows, and the grounding judge reads it against only the cited evidence. Outputs: a filing copy without citations or held sentences, a cited workpaper for the case file, and a signed report of hashes and verdicts. On the hosted route, which takes synthetic case files only, every model call gets a gateway-signed receipt that our gateway countersigns. Self-hosted, everything stays on your machine.
At a glance
- Data retention
- Nothing stored: the case file lives in memory for the request. The signed report holds hashes, citations, verdicts and receipt ids, never case text; logs carry counts only.
- What leaves the box
- Hosted: every model call goes through our gateway to the GPU serving Qwen3.8-27B, and only synthetic case files are accepted. Self-hosted on the direct route: nothing leaves the box.
- What it will not do
- Decide whether to file, or file. It never says 'compliant'. A clean result reads 'every sentence traced to the case file (review still required)'.
- Input formats
- Transactions as CSV (date and amount required; type, direction, channel, location, counterparty, memo optional), up to 400 rows; profile and alert fields as name: value pairs; notes up to 6,000 characters; your own draft up to 12,000 characters with [T3, K.name] citations.
- Typical run
- A structuring draft: a few dozen model calls, a few cents or less at the gateway list price. Checking a short draft: a few calls, a fraction of a cent. Each run shows its own measured cost.
Pick the tier for the quality you need
Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.
- In the hosted demo
Lite
numbers and citations only, no GPU
POST /sar/numbers checks a narrative you wrote with any tool: every citation must exist and every number must match or be computed from the cited rows. No drafting, no judge for the other facts.
- Models
- decosa-api SAR desk (decosa_api/verticals/sar) with the numeric grounding block (decosa_api/verticals/numeric) and the grounding module (decosa_api/verticals/grounding)
- Hardware
- Any CPU
- Quality evidence
- Correct numbers flagged, reference narratives (200 held-out synthetic cases)0 of 1,920docs/evals/sar-narrative-desk.md, part A, 26 Sep 2026
- Planted wrong numbers caught (amounts, dates, counts), reference narratives1,600 of 1,600 (1,598 before a bug fix)docs/evals/sar-narrative-desk.md, part A
- Planted wrong numbers caught in the model's own sentences340 of 356 (95.5%): amounts 133/133, dates 165/174, counts 34/40docs/evals/sar-narrative-desk.md, part D (v3)
- Latency
- no model call; milliseconds on CPU (not separately timed)
- Verification
- No proof yetNo model call, so no receipts; the numeric check is deterministic code.
- In the hosted demo
Standard
one GPU for the model (hosted demo)
Qwen3.8-27B drafts each section and judges each sentence against what it cites; the screen and the numbers are code. This is what the hosted demo runs, on synthetic case files only.
- Models
- decosa-api SAR desk (decosa_api/verticals/sar) with the numeric grounding block (decosa_api/verticals/numeric) and the grounding module (decosa_api/verticals/grounding)
- Qwen3.8-27B (NVFP4)
- Hardware
- 1x RTX PRO 6000 96 GB (measured) or 1x RTX 5090 32 GB (estimate)
- Quality evidence
- Drafted sentences traced / flagged / held (v3: 15 synthetic cases, 299 sentences)272 / 23 / 4 (1 of the 4 a false hold)docs/evals/sar-narrative-desk.md, part B
- Numbers the model wrote that failed the check (v3)0 of 533docs/evals/sar-narrative-desk.md, part B
- Citation validity (v3)0 uncited sentences; 1 unknown id of 931docs/evals/sar-narrative-desk.md, part B
- Invented facts in investigator drafts held / flagged (20 drafts)18 / 2 (all 20 caught); 0 of 172 true sentences held, 19 flagged (28 Sep re-run)docs/evals/sar-narrative-desk.md, part C
- Typology coverage: planted typology found by the screen and named in the draft (12 cases, v3)12/12 and 12/12; none named that the screen did not finddocs/evals/sar-narrative-desk.md, part B
- Hallucinated facts in traced sentences (manual read, 70 sampled)0 invented facts; 2 small unsupported details (a state, 'international')docs/evals/sar-narrative-desk.md, part E
- Latency
- measured: under a minute per full draft through the shared gateway; seconds for a no-file rationale or a short check.
- Verification
- Proof: strongEvery model call is a separate gateway call with a gateway-signed receipt; the signed report lists them all.
Also runs on
- Statements as PDFs and scansA licence-clean vision/OCR model (not chosen)not builtRead the ledger from bank statements and wire confirmations instead of CSV, with each row tied to its page. The OCR model is not chosen yet; it must be licence-clean. Hardware: 1x RTX PRO 6000 96 GB (estimate).
We host these ourselves when needed: small models get more of our own compute unless we detect a shortage, so they need no community providers.
Every model in the stack
| Model | Tiers | Params · VRAM | Verification | Details |
|---|---|---|---|---|
Red-flag screen, citation check, numeric grounding, filing copy, workpaper and signed report (no model; CPU)decosa-api SAR desk (decosa_api/verticals/sar) with the numeric grounding block (decosa_api/verticals/numeric) and the grounding module (decosa_api/verticals/grounding) 0 GBProof: partial | LiteStandard | 0 GB | Proof: partial | |
| ||||
Section drafting and the grounding judgeQwen3.8-27B (NVFP4)nvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab) 27.8B · 20 GBProof: strongIn the hosted demo | Standard | 27.8B · 20 GB | Proof: strongIn the hosted demo | |
| ||||
Bank-statement and wire-confirmation reader (alternate)A licence-clean vision/OCR model (not chosen) No proof yetSelf-host only | Alternate | n/a | No proof yetSelf-host only | |
| ||||
Tools, services and hardware
Tools
- FinCEN, Guidance on Preparing a Complete & Sufficient SAR Narrative (Nov 2003) (opens in a new tab)US government work (public domain)
The section structure of the draft and the rule of individual dates and amounts.
- FinCEN SAR Filing Instructions, version 1.0 (April 2026) (opens in a new tab)US government work (public domain)
Part V narrative checklist; supporting documentation is kept, not filed.
- FinCEN and agencies, SAR FAQs (9 Oct 2025) (opens in a new tab)US government work (public domain)
Q1 (structuring), Q4 (documenting a decision not to file is encouraged, not required).
- 31 U.S.C. 5318 (Cornell LII) (opens in a new tab)US statute (public domain)
(g)(2): SAR confidentiality.
- 31 CFR 1020.320 (eCFR) (opens in a new tab)US regulation (public domain)
(b)(3) filing deadline, (d) five-year retention, (e) confidentiality.
- FinCEN Advisory FIN-2014-A005 (funnel accounts) (opens in a new tab)US government work (public domain)
The funnel-account definition behind the screen rule.
- FinCEN Advisory FIN-2022-A002 (elder financial exploitation) (opens in a new tab)US government work (public domain)
Older adult = 60 or over; financial red flags behind the elder rule.
- scripts/eval_sar.py and docs/evals/sar-narrative-desk.mdApache-2.0
The synthetic case generator, planted-error and invented-fact tests, three model draft sets, and every result file.
- POST /sar/verify and POST /sar/numbersApache-2.0
Verify a report's signature and the hashes of the workpaper, filing copy and ledger; run the numeric check alone on any draft, with no model call.
Services
- decosa-api:8445
${DECOSA_REGISTRY}/decosa-api:<tag>GET /sar/info, /sar/samples; POST /sar/draft (SSE or JSON), /sar/numbers, /sar/verify. Keeps no case text.
- vLLM (model):8114
vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1Qwen3.8-27B NVFP4 behind our gateway (hosted) or called directly (self-host).
Hardware
- 1x RTX PRO 6000 Blackwell 96 GB Fits
Measured on our server: the hosted demo and the eval ran on this card through the shared gateway.
- 1x RTX 5090 32 GB Fits
Estimate: Qwen3.8-27B NVFP4 needs about 20 GB of weights plus KV cache; not run for this tool.
- CPU only Fits
The lite tier (numbers and citations only, POST /sar/numbers) needs no GPU.
Latency per lane
- full narrative draft (8 sections, ~22 sentences), hosted gateway route24.7 s
Measuredmeasured on our server 2026-09-26: median of 15 eval runs (11.7-61 s) with the gateway shared with other workloads; 10-11 s when it was quiet
- no-file rationale (4 sections), hosted gateway route12.0 s
Measuredmeasured on our server 2026-09-26: 3 eval runs, 11.7-13.7 s
- check an investigator's draft (10 sentences)3.2 s
Measuredmeasured on our server 2026-09-26: the planted-errors sample, 2.2-10 s depending on gateway load
- numeric check only (POST /sar/numbers)n/a
Measuredno model call; milliseconds on CPU (not separately timed)
Notes
- Numbers are checked in code, not by the model. On 200 held-out synthetic cases the checker flagged none of the 1,920 correct numbers in reference narratives and caught all 1,600 planted wrong amounts, dates and counts (1,598 before a bug fix, see the eval). Planted into the model's own sentences it caught 340 of 356 (95.5%): every amount, but 165 of 174 dates and 34 of 40 counts.
- In the final eval set (15 synthetic cases, 299 drafted sentences) the model's 533 numbers all checked out; 272 sentences were traced, 23 flagged as partly supported, 4 held. Of the 4 held, one was a false hold by the judge.
- Invented facts: of 20 investigator drafts with one invented sentence each, 18 were held and 2 flagged, so all 20 were caught; of 172 true sentences none was held and 19 were flagged (re-run 28 Sep 2026 after a change that cut flags on true sentences from 27 to 19; the 26 Sep run held 19 and flagged 1 invented sentence, held 1 true sentence and flagged 32).
- Typologies: the screen found the planted typology in 160 of 160 held-out cases, but also flagged structuring on 5 of 40 clean cash-business cases, and funnel cases also trip structuring (13 of 40). The thresholds are ours, set on our own generator.
- The judge is strict: it flags small inferences (a state read off a branch name, 'domestic', 'international') as partly supported. In a manual read of 70 traced sentences, none held an invented fact and 2 carried such a small unsupported detail.
- Everything here is measured on synthetic case files written by the agent that built the checker. It has not been run on real bank data.
Run this exact stack on your machine
Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.
# Assemble the Decosa SAR narrative desk on this machine
You are setting up a drafting aid for BSA/AML investigators. From an alert's case file (a transactions CSV, the customer
profile and alert fields, the investigator's notes) it drafts a suspicious-activity narrative in FinCEN's structure (who,
what, when, where, why, how), or a no-file rationale, or checks a draft the investigator wrote. Every sentence cites the
rows, fields or notes it rests on; every number is recomputed in code from the cited rows; the other facts are judged
against only the cited evidence. It returns a filing copy, a cited workpaper and a JSON report signed by this box's own
key. Work step by step, show me each command before you run anything with `sudo`, and stop to ask if a check fails.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/sar-narrative-desk.zip (4 KB, 15 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py sar-narrative-desk` (the api image carries the same bundle under /app/rehearsal/sar-narrative-desk/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py sar-narrative-desk --bundle sar-narrative-desk.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the run holds sentences back", "the wrong total is held as a number mismatch", "the held total names the figure the rows give"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## 0. Ground rules and licences
- Model: Qwen3.8-27B (Apache-2.0), used for two jobs: drafting each section with citations, and the grounding judge
(is the sentence backed by what it cites). The screen, the numeric check and the report are decosa-api (AGPL-3.0-or-later)
and need no GPU of their own.
- SAR information is confidential (31 U.S.C. 5318(g)(2); 31 CFR 1020.320(e)). Case files stay on this machine. Bind
every port to 127.0.0.1. The service keeps no case text: nothing is written to disk and logs carry counts only. Keep
it that way; do not add request logging.
- Be honest about what it is: a drafting aid. The investigator decides whether to file, edits the narrative and files
it. It is never a compliance determination, and the red-flag screen is decision support with our own thresholds.
## 1. Check the machine
1. `nvidia-smi`: one GPU with at least 32 GB (Qwen3.8-27B NVFP4 needs about 20 GB of weights plus KV cache; an RTX PRO
6000 96 GB is what we measured on; an RTX 5090 32 GB should fit but we have not run this tool on one). Driver
570 or newer. Blackwell cards run NVFP4; on older cards use `Qwen/Qwen3.8-27B-FP8`.
2. `docker --version` and `docker compose version`. If Docker or the NVIDIA container toolkit is missing, install them
from the official Docker and NVIDIA repositories after asking me, then run
`docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
3. Disk: about 30 GB free.
## 2. Images and weights
- `${DECOSA_REGISTRY}/decosa-api:<tag>` (**publishing soon**). If the pull fails, build from source:
`git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host>` (access required), check out the newest release tag that
contains `decosa_api/verticals/sar/` and `decosa_api/verticals/numeric/` (`main` until one does), and build
`docker/api/Dockerfile`.
- `vllm/vllm-openai:v0.29.0` for the model; weights `nvidia/Qwen3.8-27B-NVFP4` (revision
`482ca0f3832238542f8f5295dde86b5f22711d80`), or `Qwen/Qwen3.8-27B-FP8` on a card without NVFP4.
## 3. docker-compose.yml
Write this in `~/decosa/sar/`:
```yaml
services:
llm:
image: vllm/vllm-openai:v0.29.0
command: ["--model", "nvidia/Qwen3.8-27B-NVFP4", "--served-model-name", "qwen3.8-27b", "--max-model-len", "32768",
"--enable-prefix-caching"]
ports: ["127.0.0.1:8114:8000"]
volumes: ["~/.cache/huggingface:/root/.cache/huggingface"]
deploy: { resources: { reservations: { devices: [{ driver: nvidia, count: 1, capabilities: [gpu] }] } } }
healthcheck: { test: ["CMD", "curl", "-fs", "http://localhost:8000/v1/models"], interval: 30s, retries: 20 }
api:
image: ${DECOSA_REGISTRY}/decosa-api:<tag>
ports: ["127.0.0.1:8445:8445"]
environment:
DECOSA_HOST: 0.0.0.0
DECOSA_PORT: "8445"
DECOSA_DATA_DIR: /data
DECOSA_LLM_ROUTE: direct
DECOSA_LLM_URL: http://llm:8000/v1
DECOSA_LLM_MODEL: qwen3.8-27b
DECOSA_SAR_SYNTHETIC_ONLY: "0"
DECOSA_SAR_MAX_CONCURRENT: "3"
DECOSA_SAR_WORKERS: "6"
DECOSA_BUDGET_LLM_TOKENS: "60000"
volumes: ["decosa-data:/data"]
depends_on: { llm: { condition: service_healthy } }
healthcheck: { test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://localhost:8445/sar/info', timeout=4)"], interval: 30s, retries: 10 }
volumes:
decosa-data:
```
`DECOSA_SAR_SYNTHETIC_ONLY: "0"` lets this box take real case files; the hosted service refuses anything not marked
`"synthetic": true`. The api keeps its state (keys, receipts, this box's signing key) in the named volume `decosa-data`,
not in a host folder: the image runs as an unprivileged user (uid 10001), and a host folder that Docker creates is owned
by root, which stops the api with `PermissionError: [Errno 13] Permission denied: '/data/keys.sqlite'`. Then start
everything: `docker compose up -d`.
Prefix caching matters: every section call and every sentence check of one case sends the same case file first.
On the first start the api creates this box's Ed25519 key in the `decosa-data` volume (`/data/attest/`, mode 0600).
Back it up with `docker compose cp api:/data/attest ./attest-backup` and keep that copy private. Never print it. Every
model call on the direct route gets a receipt signed with that key (status `attested`): an attestation by me, the
operator, not a proof of computation. Never set `DECOSA_LLM_ROUTE=gateway` on this box: that sends case data to the
hosted Decosa API.
## 4. Smoke test
1. `curl -s localhost:8445/sar/info | jq '{modes: (.modes | keys), synthetic_only}'` shows `draft`, `no_file`,
`check` and `synthetic_only: false`.
2. Token: `T=$(curl -s -XPOST localhost:8445/demo/session -H 'content-type: application/json' -d '{"vertical":"sar-narrative-desk"}' | jq -r .token)`.
3. `curl -s localhost:8445/sar/samples > samples.json`, then
`jq '.[] | select(.id=="planted-errors") | {transactions, kyc, alert, notes, mode, narrative, synthetic}' samples.json > planted.json` and
`curl -s -XPOST localhost:8445/sar/draft -H "authorization: Bearer $T" -H 'content-type: application/json' -d @planted.json > res.json`.
Expect `status: "held"` with three held sentences: the wrong total and the shifted date as `number_mismatch`, each
with the right figure in `numeric.mentions[].detail`, and the invented $25,000.00 wire. The other sentences are
mostly `checked`.
4. Draft one: take `.[0]` (structuring) the same way (mode `draft`). Expect eight sections, `screen.indicated` =
`["structuring"]`, and most sentences `checked`. A few `flagged` or `held` sentences are normal; read why.
5. `jq '{report, narrative_md, narrative_plain}' res.json | curl -s -XPOST localhost:8445/sar/verify -H 'content-type: application/json' -d @-`
must show `valid_signature`, `signed_by_this_server` and both `..._matches` true. Change `status` in the report and
verify again: it must fail.
6. Stream one with `-H 'accept: text/event-stream' -N`: `ready`, then `section`, `receipt` and `sentence` events, then
`report`, `budget` and `done`.
7. Time it and tell me what you measure. The check sample took about 10 s and a full draft 20-80 s on our shared RTX
PRO 6000 through the gateway; the direct route is usually faster.
## 5. Point the app at the local API
Set `NEXT_PUBLIC_DECOSA_API=http://127.0.0.1:8445` in the site's `.env.local`, or call `POST /sar/draft` from your case
management tool and attach `narrative_md` (the workpaper) and the signed report to the case. `POST /sar/numbers` runs
the numeric check alone on a draft from any other tool, with no model call. Contract: `API_CONTRACT.md`, section "SAR
narrative desk".
Off by default. Joining serves other people's requests on this GPU; never do it on a box that holds SAR case data. If I
ask for it, follow the provider guide at `/provide` on the site, and do not enable it without my explicit yes.Rules and regulations it checks againstDated, linked to the primary source; not legal advice
Regulation watch
Loading the watch status…
10 laws, rules and guidance pages cited; 10 watched nightly at the primary source. A change marks this page for a human re-check; nothing is edited automatically. What we cite and how it is watched
Technical detailsModels, where it runs, labels
In short
Last reviewed
- What it is
- A suspicious-activity narrative drafted from the case file, with every sentence citing its rows and every number recomputed in code.
- Who it's for
- Teams in finance and insurance and compliance and trust.
- Where it runs
- Self-host for real cases (SAR confidentiality); the hosted demo takes synthetic case files only
- Key numbers
- 1,598 of 1,600 (99.9%) Planted wrong numbers caught in reference narratives, first run (no model) (test split, n = 1600)
- 340 of 356 (95.5%) Wrong numbers planted in the model's own sentences caught (v3) (test split, n = 356)
- 18 of 20 Invented sentences held in investigator drafts (test split, n = 20)
- 24.7 s Median end-to-end run, hosted (QA sweep 2026-09-26)
- Models
- Qwen3.8-27B drafts and judges; the numbers and the red-flag screen are plain code
- Where
- Self-host for real cases (SAR confidentiality); the hosted demo takes synthetic case files only
- Checks
- Receipt per model call; every number recomputed from the cited rows; signed report over hashes and verdicts
- Industry
- Finance and insurance · Compliance and trust
- Input
- Text and documents
- Runs
- Self-host
- Output
- Notes, reports and drafts · Signed record or verdict
- Data
- Confidential business data · Personal data
- Hardware
- 1× 96 GB GPU
- Licence
- Permissive (Apache-2.0, MIT)
- Runs in
- Self-host
- Built from
- Numeric grounding · Grounding · Signed record
Questions people ask
How are the numbers checked?
In code, not by the model. Every amount, date, count, span and percentage is recomputed from the rows the sentence cites, and a wrong one is held back from the filing copy with the figure the rows do support.
How well does the number check work?
On 200 held-out synthetic cases it caught 1,598 of 1,600 planted wrong numbers on the first run (1,600 after a fix) and flagged none of 1,920 correct ones. Planted into the model's own sentences it caught 340 of 356. All of this is on our own synthetic generator, not real bank data.
Can I send real case files to the hosted demo?
No. SAR information is confidential (31 U.S.C. 5318(g)(2); 31 CFR 1020.320(e)), so it is self-host first and the hosted demo refuses any case file not marked synthetic. Self-hosted on the direct route, nothing leaves the box.
What structure does the draft follow?
FinCEN's narrative guidance: introduction; who, what, when, where, why and how; conclusion. The filing copy drops the citations, and the cited workpaper stays in the case file, since supporting documentation is not filed with the SAR.
Does it decide whether to file?
No. It never decides to file or files, and never says 'compliant'. It can draft an optional no-file rationale, which FinCEN encourages but does not require, and it can check a narrative you wrote.
Ask a question or leave feedbackWe read every message and publish useful answers
Ask about SAR narrative desk
We read every message. Questions, comments and our answers show here once we have reviewed and approved them.
Loading questions…