Build a model-risk evidence pack
An evidence pack showing whether an LLM you rely on still behaves as validated, from which your validators can recompute every number.
Built on: Endpoint audit, Typed judgment, Signed record
Loading the tool…
Use it your way
Use it from your codeThe hosted API with your key, and prompts to paste into a coding agent
Get an API key
- Call the model-risk evidence pack API from your own code in minutes.
- Every model answer carries a signed receipt.
- Nothing to install; we run the models.
Run it yourself, on request
- The same open models and app, on 1× RTX PRO 6000 (96 GB) or 1× RTX 5090 (32 GB) for the grader; the pack runner is CPU; the system under test runs wherever it runs.
- Data never leaves your machines, and there are no Decosa charges.
- One prompt for Claude Code or Codex assembles the whole stack.
- Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Build with it
Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.
- Base URL
- https://api.decosa.ai
- Auth
Authorization: Bearer $DECOSA_API_KEY(or a demo session token)- Tool id
- model-risk-pack
Use the hosted API
# Decosa model-risk evidence pack: use the hosted API
You are wiring Decosa's model-risk evidence pack into this project. It re-runs a fixed test suite against an LLM system
we rely on (the hosted demo systems, or our own OpenAI-compatible endpoint): identity prompts, graded cases, stability
against a validated run, counterfactual fairness pairs. Flagged lanes are re-run before any alert. It returns a signed
pack, a Markdown binder and a hash-chained record our validators can recompute. Use only what is listed below. If you
need something else, stop and ask me.
- Base URL: `https://api.decosa.ai`
- Health check: `GET https://api.decosa.ai/healthz`.
- The pack is evidence, not a validation opinion and not legal advice. It detects only what its suite probes.
- The hosted demo is for synthetic cases. Real complaints or applications belong on a self-hosted box.
## Auth: API key (or a demo session)
1. Preferred: an API key (`dk_…`) from "Get an API key" on the tool page. Keep it in `DECOSA_API_KEY`, never in
code. Send `Authorization: Bearer $DECOSA_API_KEY`.
2. Without a key: `POST https://api.decosa.ai/demo/session` with `{"vertical": "model-risk-pack"}` returns `{"token", "expires_at", "budget"}`.
20,000 generated tokens per session: about three full packs.
3. A run needs at least 4,500 generated tokens left (402 otherwise). One run at a time per token.
## The run
`POST /mrm/runs` (token). Body, one of:
```json
{"target": "vendor-biased"}
{"endpoint": {"base_url": "https://api.example.com/v1", "model": "our-model", "api_key": "…", "black_box": true},
"suite": "fernhill-v1", "baseline": { "…the first signed pack…": "" }, "history": [], "system": {"owner": "Model risk", "version": "2026-Q3"}}
```
- `target`: a demo system from `GET /mrm/targets` (`lender`, `vendor`, and the vendor after injected changes). Demo
targets use their stored, signed baseline.
- `endpoint`: public https only; the key is used for this run and never logged or stored. `black_box: true` sends only
the case text (the endpoint applies its own prompt); otherwise the suite's prompts are sent as the system message.
- `suite`: `fernhill-v1` (synthetic, `GET /mrm/suites/fernhill-v1`) or our own suite object in the same schema.
- `baseline`: our first pack for this system and suite. Without it the pack records the validated behaviour.
- JSON by default: `{pack, record, markdown, cost_usd, budget}`. With `Accept: text/event-stream` (or `"stream": true`):
`ready`, receipt and `probe` events as each call lands, `lane` events, `recheck` (if a lane was flagged), `verdict`,
`pack`, `record`, `markdown`, `budget`, `done`.
- `pack.verdict.status`: `pass`, `watch` (flagged once, not reproduced), `alert` (flagged and reproduced) or
`inconclusive` (too many failed calls). `pack.flags` says which lane and why; the rules are `pack.policy`.
Other endpoints (no token): `GET /mrm/info`, `POST /mrm/verify` `{pack, record, suite?, baseline?}` →
`{ok, pack, record, recompute}`, `POST /mrm/render` `{pack}` → the Markdown binder, `POST /mrm/trend` `{packs}`,
`GET /attest/signing-key`, `GET /receipts/{id}`.
## Example: a weekly check of our vendor endpoint (Python, `pip install httpx`)
```python
import httpx, json, os, pathlib
API = "https://api.decosa.ai"
H = {"Authorization": f"Bearer {os.environ['DECOSA_API_KEY']}"}
d = pathlib.Path("mrm-packs"); d.mkdir(exist_ok=True)
body = {"endpoint": {"base_url": os.environ["VENDOR_URL"], "model": "triage-v3", "api_key": os.environ["VENDOR_KEY"], "black_box": True},
"suite": json.load(open("our-suite.json"))}
if (d / "baseline.json").exists():
body["baseline"] = json.load(open(d / "baseline.json"))
r = httpx.post(f"{API}/mrm/runs", json=body, headers=H, timeout=900)
r.raise_for_status()
res = r.json()
pack = res["pack"]
if "baseline" not in body:
json.dump(pack, open(d / "baseline.json", "w"))
json.dump(pack, open(d / f"{pack['id']}.json", "w")); json.dump(res["record"], open(d / f"{pack['id']}.record.json", "w"))
open(d / f"{pack['id']}.md", "w").write(res["markdown"])
print(pack["verdict"]["status"], pack["verdict"]["reasons"])
```
## Keep and verify
Keep the pack, the record and the suite together. `POST https://api.decosa.ai/mrm/verify` checks the pack signature and every
record entry (text hashes, links, embedded receipts), then re-parses and re-grades every output from the record's texts
and recomputes each number and the verdict. Pin the server's key from `/attest/signing-key`.
## Honest limits
- A problem the suite does not probe (another product, another attribute, a ZIP code it does not use) is not detected.
- Drafts are graded by an open model (typed judgments); the grade is repeatable, not infallible.
- For a black box, identity cannot tell a model swap from a prompt or settings change; it shows that behaviour changed.
- Latency is measured on a shared GPU through the gateway.
Run it yourself (containers)
On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.
# Decosa model-risk evidence pack: run it yourself (containers)
You are setting up the Decosa model-risk evidence pack on this machine, so our test cases and the systems we test stay
inside our network. It re-runs a signed test suite against an LLM system we use, re-checks anything it flags, and signs
a pack and a hash-chained record with this box's key. Nothing is sent to Decosa's hosted API.
Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.
Ask me before any command that needs sudo, and show me the command first.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/model-risk-pack.zip (3 KB, 10 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py model-risk-pack` (the api image carries the same bundle under /app/rehearsal/model-risk-pack/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py model-risk-pack --bundle model-risk-pack.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the run finishes without errors", "all 14 calls were made (10 identity probes and 4 cases)", "no call failed"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
instructions (docs.docker.com/engine/install), plus the NVIDIA container toolkit; check
`docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
`mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
Keep the `llm` service (Qwen3.8-27B on vLLM: the grader) and the `api` service. For `api` set
`DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b`,
`DECOSA_MRM_ALLOW_PRIVATE=1` (so we can test endpoints on our own network) and bind every port to 127.0.0.1.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check (the first start
downloads about 20 GB of weights).
4. Check: `curl -fsS http://127.0.0.1:<PORT>/mrm/info` lists the lanes, the policy and the grader;
`GET /attest/signing-key` shows this box's public key. Show me the key: validators pin it.
5. Smoke test: get a token with `POST /demo/session {"vertical":"model-risk-pack"}`, then `POST /mrm/runs {"target": "lender"}`.
Expect about 80 calls, triage 14-15 of 15, drafts 4-5 of 5 and a verdict. Then `POST /mrm/verify` with the pack
and the record: `ok`, `record.ok` and `recompute.ok` must all be true.
6. Our own system: `POST /mrm/runs {"endpoint": {"base_url": "http://…/v1", "model": "…", "black_box": true}, "suite": <our suite>}`.
Keep the first pack as the baseline and send it as `baseline` every time after. Schedule `scripts/mrm_monitor.py`
weekly (cron or a systemd timer); it exits 2 on an alert.
7. Report back: the public key and key id, the smoke-test verdict, and how long the run took.
Off by default. Never on a box that holds confidential test cases. If I ask for it later, follow the Provide page.
No NVIDIA GPU? This tool also runs entirely on an Apple Silicon Mac (MLX, 32 GB of unified memory or more): use https://decosa.ai/prompts/model-risk-pack-mac.md instead.
Run it on your own hardwareWhat it needs, and the prompt that sets it up
Run it on your own GPU
Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.
Hardware check
Check your own hardware- CPU only, 64 GB RAMDoesn't fit
Qwen3.8-27B (NVFP4) needs a GPU.
- GeForce RTX 4090standard tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)
- GeForce RTX 5090standard tierRuns
The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
- 2x GeForce RTX 5090standard tierRuns
The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size).
- L40Sstandard tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.
- H100 80 GB (SXM)standard tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.
- RTX PRO 6000 Blackwell 96 GBstandard tierRuns
The standard tier fits (57.6 of 96 GB).
- 2x RTX PRO 6000 Blackwell 96 GBstandard tierRuns
The standard tier fits (57.6 of 192 GB).
- Apple M3 Ultra (Mac Studio), 96 GBstandard tierRuns
The standard tier fits (32 of 96 GB).
- Apple M5 Max, 64 GBstandard tierRuns
The standard tier fits (32 of 64 GB).
Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.
On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.
- 1
Check the GPU, Docker and the NVIDIA Container Toolkit
The driver must see the GPU, and Docker must be able to pass it into a container.
nvidia-smi docker compose version docker run --rm --gpus all ubuntu nvidia-smi
- 2
Fetch the compose file
One file describes the API and the language model as services.
mkdir -p ~/decosa && cd ~/decosa curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml - 3
Pull and start
The first start downloads pinned model weights, tens of gigabytes.
docker compose pull docker compose up -d
- 4
Check health
Wait until the API reports ok with the language model loaded. Then point your app at the local base URL.
curl -fsS http://localhost:<PORT>/healthz # {"ok": true, "llm": true, ...} curl -fsS -X POST http://localhost:<PORT>/demo/session \ -H 'Content-Type: application/json' -d '{"vertical":"model-risk-pack"}'
Set up with a coding agent, rehearse on mock data, then go private
- Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
- Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's
expected.json. Every check must print PASS. - Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
docker compose exec api python scripts/rehearse.py model-risk-pack
Download the mock-data bundle (3 KB, 10 checks)expected.json
A cut-down validation suite (four synthetic member complaints for the fictional Fernhill Valley Credit Union, with the expected queue and flags) run against the documented demo deployment, plus the ten identity probes. The planted cases are a servicemember asking for SCRA protection (must be urgent) and a mortgage applicant describing discrimination (must be flagged). The signed pack and its record must verify and recompute, and a pack with one observed answer changed must not.
What the rehearsal checks
- the run finishes without errors
- all 14 calls were made (10 identity probes and 4 cases)
- no call failed
- at least 3 of the 4 cases are triaged correctly
- the planted SCRA request from a servicemember is marked urgent
- the planted discrimination allegation is flagged
- the pack verdict is pass or watch
- the signed pack, its record and the recomputation all verify
- a pack with one observed answer changed no longer verifies
- every model call has a signed receipt
Licence: Synthetic: cases written for the Fernhill suite on 25 Sep 2026; Fernhill Valley Credit Union is fictional and names and complaints are invented. Part of decosa-api, AGPL-3.0-or-later.
Prompt for your coding agent
# Decosa model-risk evidence pack: run it yourself (containers)
You are setting up the Decosa model-risk evidence pack on this machine, so our test cases and the systems we test stay
inside our network. It re-runs a signed test suite against an LLM system we use, re-checks anything it flags, and signs
a pack and a hash-chained record with this box's key. Nothing is sent to Decosa's hosted API.
Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.
Ask me before any command that needs sudo, and show me the command first.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/model-risk-pack.zip (3 KB, 10 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py model-risk-pack` (the api image carries the same bundle under /app/rehearsal/model-risk-pack/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py model-risk-pack --bundle model-risk-pack.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the run finishes without errors", "all 14 calls were made (10 identity probes and 4 cases)", "no call failed"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
instructions (docs.docker.com/engine/install), plus the NVIDIA container toolkit; check
`docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
`mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
Keep the `llm` service (Qwen3.8-27B on vLLM: the grader) and the `api` service. For `api` set
`DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b`,
`DECOSA_MRM_ALLOW_PRIVATE=1` (so we can test endpoints on our own network) and bind every port to 127.0.0.1.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check (the first start
downloads about 20 GB of weights).
4. Check: `curl -fsS http://127.0.0.1:<PORT>/mrm/info` lists the lanes, the policy and the grader;
`GET /attest/signing-key` shows this box's public key. Show me the key: validators pin it.
5. Smoke test: get a token with `POST /demo/session {"vertical":"model-risk-pack"}`, then `POST /mrm/runs {"target": "lender"}`.
Expect about 80 calls, triage 14-15 of 15, drafts 4-5 of 5 and a verdict. Then `POST /mrm/verify` with the pack
and the record: `ok`, `record.ok` and `recompute.ok` must all be true.
6. Our own system: `POST /mrm/runs {"endpoint": {"base_url": "http://…/v1", "model": "…", "black_box": true}, "suite": <our suite>}`.
Keep the first pack as the baseline and send it as `baseline` every time after. Schedule `scripts/mrm_monitor.py`
weekly (cron or a systemd timer); it exits 2 on an alert.
7. Report back: the public key and key id, the smoke-test verdict, and how long the run took.
Off by default. Never on a box that holds confidential test cases. If I ask for it later, follow the Provide page.
No NVIDIA GPU? This tool also runs entirely on an Apple Silicon Mac (MLX, 32 GB of unified memory or more): use https://decosa.ai/prompts/model-risk-pack-mac.md instead.
Help me customise for my hardware
Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.
GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page
RunsModel-risk evidence pack on GeForce RTX 5090: use the Standard · hosted grader and demo systems (Qwen3.8-27B) tier
The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
What this tool's stack says about this hardware:
- 1x RTX 5090 32 GB (fits): Qwen3.8-27B NVFP4 needs about 20 GB of weights plus KV cache. Estimate: not run for this vertical.
Standard · hosted grader and demo systems (Qwen3.8-27B): what changesuses estimates
- Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
Memory per component
- Pack runner: decosa-api model-risk module (decosa_api/verticals/mrm). CPU. Runs on CPU (vram_gb 0 in stack.json).
- Fixed grader: Qwen3.8-27B (NVFP4). ~57.6 GB (at least ~28 GB), weights 21.4 GB (from stack.json). Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)
Expected speed
Not measured.
Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.
Setup prompt for this hardware
The self-host prompt for Model-risk evidence pack, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.
# Set up Model-risk evidence pack on my hardware Fetch https://decosa.ai/prompts/model-risk-pack-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied. ## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=model-risk-pack) Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4). Quality tier: Standard · hosted grader and demo systems (Qwen3.8-27B) (standard). Fit check: runs with changes, about 28 GB of 32 GB used; some memory numbers are estimates, not measurements. First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything. Use these components (the setup below describes the standard tier; change it to match): - Pack runner: decosa-api model-risk module (decosa_api/verticals/mrm), CPU - Fixed grader: Qwen3.8-27B (NVFP4) (nvidia/Qwen3.8-27B-NVFP4), 57.6 GB. Change: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions. GPU placement (set each service's device and its vLLM --gpu-memory-utilization to about the share shown): - GPU 0: Qwen3.8-27B (NVFP4) ~28 GB (88%); about 4 GB left During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed. The stack's own component list and compose layout: https://decosa.ai/prompts/model-risk-pack-assemble.md
Or on a Mac Studio
No NVIDIA GPU needed: every model this tool uses runs natively on Apple Silicon through MLX. Any M-series Mac with 32 GB of unified memory or more. Measured speeds and what runs where
From a checkout of decosa-api, one command sets up the models and the API: scripts/mac/setup.sh
Mac prompt for your coding agent
# Decosa Model-risk evidence pack: run it on this Mac (Apple Silicon, no NVIDIA GPU) You are setting up the Decosa Model-risk evidence pack on this Mac, natively on Apple Silicon. The models run on the Mac's GPU through MLX and decosa-api runs from a git checkout with `uv`. Docker is not used for the models, because Docker on macOS cannot reach the GPU. Nothing is sent to Decosa's hosted API. Every model this tool needs runs on the Mac. It needs 32 GB of unified memory or more. Ask me before any command that needs sudo or installs software with Homebrew, and show me the command first. Never stop or kill a process this setup did not start; if a port is taken, pick another one. ## Step 0: set up with a coding agent, rehearse on mock data, then go private This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works. Work in this order: 1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to "test with something realistic". 2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool, https://decosa.ai/samples/model-risk-pack.zip (3 KB, 10 checks, synthetic or openly licensed: see `licence` in expected.json), show me what is in it, and run the rehearsal against the local API: `.venv/bin/python scripts/rehearse.py model-risk-pack` in the decosa-api checkout (the key comes from ~/.decosa-mac/api.key). It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the run finishes without errors", "all 14 calls were made (10 identity probes and 4 cases)", "no call failed"). Show me the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json` to make a check pass. 3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this machine. For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent can read. Switch to your own data only after the rehearsal has passed and the agent's work is done. ## What runs where | Part | On an NVIDIA GPU | On this Mac | Status | |---|---|---|---| | Pack runner: suite, calls, parsing, grading rules, stability, fairness, re-check, model card, signed pack and record (no model; CPU) | Python on CPU | The same Python module, run with uv | Runs, measured | | Fixed grader (typed judgments) and the hosted system under test | NVFP4 on vLLM 0.29 (Blackwell) | MLX 4-bit (EigenLabs/Qwen3.8-27B-4bit) on mlx_lm.server 0.31.3; oMLX 0.6.1 with MTP as an option | Runs, measured | ## Steps 1. Check the machine: `uname -m` must print `arm64` (an M-series chip; Intel Macs cannot run MLX), and `sysctl -n hw.memsize` should be at least 32 GB for this tool. Check about 30 GB of free disk with `df -h ~`. Show me the chip (`sysctl -n machdep.cpu.brand_string`) and the memory. 2. Tools: `uv --version`. If it is missing, ask me, then `brew install uv`. 3. Code: `git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host> ~/decosa-api` (access required) and `cd ~/decosa-api`. Check that `scripts/mac/setup.sh` exists; if it does not, the checkout is too old: stop and tell me. 4. Start everything with one command: `scripts/mac/setup.sh`. It creates `.venv` (decosa-api) and `.venv-mac` (MLX, mlx-lm, mlx-audio), downloads the weights with the Hugging Face CLI (about 16 GB for the language model), starts the model servers and decosa-api on 127.0.0.1, and mints a local API key into `~/.decosa-mac/api.key` (mode 0600). The first run takes a while because of the downloads; later runs reuse them. If a download fails with 401 or 403, ask me for a Hugging Face token and set `HF_TOKEN`. 5. Check health: `scripts/mac/setup.sh status` shows each server, and `curl -fsS http://127.0.0.1:8445/healthz` must report `"llm": true`. `curl -fsS http://127.0.0.1:8445/attest/signing-key` shows this Mac's public key: show it to me, because it is what others pin to check the receipts and records this Mac signs. 6. Smoke test: `.venv/bin/python scripts/mac/bench_usecases.py model-risk-pack`. It runs the tool's own sample end to end against the local API with the local key and prints `ok`, the wall time, the model calls and the receipts. `ok=True` is the pass condition. If it fails, read `~/.decosa-mac/logs/*.log` and tell me what you found. 7. Point the app at it: the API is `http://127.0.0.1:8445` with `Authorization: Bearer $(cat ~/.decosa-mac/api.key)`, the same routes as the hosted API. To stop everything: `scripts/mac/setup.sh stop`. 8. Report back: the chip and memory, the public key, the smoke-test result and its time, and the output of `scripts/mac/setup.sh status`. ## Good to know - Receipts: every model call is signed with this Mac's own Ed25519 key and names the exact MLX weights (`qwen3.8-27b-mlx-4bit` with a hash of the downloaded files). There is no gateway countersignature on a self-hosted Mac. - The weights are a 4-bit MLX build of the same open models, not the NVFP4 build the hosted route and the published evals use. Expect small differences in wording and scores. - Faster drafting: `scripts/mac/setup.sh stop && scripts/mac/setup.sh --engine omlx` serves the model with oMLX and multi-token prediction (about 2x faster for a single long answer, no faster for many parallel calls; typed judgments then use sampling because oMLX returns no log-probabilities). - Measured speeds for a Mac Studio M3 Ultra and the memory each tool needs: https://decosa.ai/mac. Full details: `docs/self-host-mac.md` in the checkout.
The proof
How we tested itEval results and end-to-end checks, hosted and self-hosted, with dates
Verified end to end
Hosted: verified 25 Sep 2026 · measured 25 Sep 2026: · p50 62 s · ~$0.013 per run · 82 receipts
Loading the nightly status…
Self-host: verified 25 Sep 2026 · Fresh clone of the pre-release branch, api image built from docker/api/Dockerfile, compose from the assemble prompt (llm service dropped, api on host network pointed at the running Qwen3.8-27B vLLM, named volume).
Measured cost to run: about $0.013 per pack (hosted, 25 Sep 2026). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.
Verified on 2026-09-25: image builds, service starts healthy, sample passes end to end against local model servers equivalent to the documented ones; model-server startup itself not re-verified. Lender pack PASS (15/15, 5/5, 82 attested receipts, 12 s); /mrm/verify ok, and a one-word edit in the record fails at that entry; key minting and two mrm_monitor.py runs (baseline, then trend) worked; logs held no case text. Torn down afterwards.
Known limits (5)
- Detects only what the suite probes: other products, attributes or ZIP codes are not covered.
- Sampling or settings changes that do not change decisions are not detected (0 of 4 in the eval).
- The drafts grader checks that listed reasons appear, not how specific they are.
- Hosted demo sessions allow about three packs (20,000 generated tokens); use an API key for more.
- Scheduled monitoring is a client script (cron or systemd), not a hosted scheduler.
How it's builtThe steps, the models and what each one checks
Get an API key
- Call the model-risk evidence pack API from your own code in minutes.
- Every model answer carries a signed receipt.
- Nothing to install; we run the models.
Run it yourself, on request
- The same open models and app, on 1× RTX PRO 6000 (96 GB) or 1× RTX 5090 (32 GB) for the grader; the pack runner is CPU; the system under test runs wherever it runs.
- Data never leaves your machines, and there are no Decosa charges.
- One prompt for Claude Code or Codex assembles the whole stack.
- Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Signed evidence that an LLM your bank, insurer or lending team relies on still behaves as validated, with a record your validators can recompute.
Re-runs a fixed test suite against an LLM system you use: your own deployment of an open model, a vendor's black-box API, or any OpenAI-compatible endpoint. Four lanes: identity (the endpoint auditor's golden prompts against the validated run), graded canaries (cases scored exactly, and drafts graded by typed judgments on a fixed open grader), stability (the same inputs as the validated run) and fairness (counterfactual pairs that change only a name, ZIP code or age). A flagged lane is run again, independently, before the pack says alert. You get a signed pack with a model card and a trend line, a Markdown binder, and a hash-chained record of every output with its receipt, from which /mrm/verify recomputes every number. The demo is a fictional credit union's complaint triage and adverse-action reasons.
- Deployment
- Hosted or self-host
- Regulatory
- Checked 25 Sep 2026; not legal advice. US banks: SR 26-2 (Fed, OCC, FDIC, 17 Apr 2026) replaced SR 11-7 and OCC 2011-12, and its footnote 3 puts generative and agentic AI models outside its scope, leaving their controls to each bank's own risk management and governance. This pack is evidence for that governance; it is not a validation opinion and does not make anyone compliant. Lenders: ECOA / Regulation B (12 CFR 1002.9) requires specific principal reasons for adverse action; the CFPB withdrew Circulars 2022-03 and 2023-03 on 12 May 2025, but the rule stands. Insurers: the NAIC AI model bulletin is adopted in 25 jurisdictions (NAIC map, 1 Apr 2026), insurers stay responsible for vendor AI systems, and a 12-state AI Systems Evaluation Tool pilot ran March to September 2026; NYDFS Circular Letter 7 (2024) asks for testing for unfair discrimination before and after deployment. EU: credit scoring of natural persons is high-risk under the AI Act (Annex III 5(b)); after Regulation (EU) 2026/1744 those obligations apply from 2 Dec 2027. The fairness probes change one attribute at a time among the values in the suite; they do not measure disparate impact on real applicants. Test cases often hold customer data: keep real cases on a self-hosted box. Model licences: Apache-2.0 (Qwen3.8-27B, Qwen3-1.7B).
Text description
A signed test suite and the validated baseline pack go to the pack runner, which needs no GPU. The runner calls the system under test: the hosted open model through our gateway, a vendor's black-box API, or your own endpoint. It computes four lanes: identity, graded canaries, stability against the baseline and counterfactual fairness flips. Drafts are graded by typed judgments on a fixed open grader, Qwen3.8-27B. A flagged lane runs a second time before an alert. Outputs: a signed evidence pack with a model card and trend, a Markdown binder, and a hash-chained record of every output with its receipt, which a validator re-checks and recomputes. Each call's receipt is countersigned by the gateway on the hosted route. In self-host mode the runner, grader and test cases stay on your machine.
At a glance
- Typical run cost
- A few cents or less per pack at the gateway list price; more with a re-check. Each run shows its own measured cost.
- Data retention
- Hosted: packs for the demo systems (synthetic) are stored; packs for your endpoint or suite are never stored, they go back to you. Self-host: nothing leaves the box.
- What leaves the box
- Hosted: your suite's case texts go to the system under test and, for drafts, to the hosted grader. Self-host: only calls to the endpoint you test.
- Inputs
- A suite of up to 40 classify cases, 10 fairness bases with up to 8 variants, 10 drafts (200 KB); any OpenAI-compatible endpoint, https and public on the hosted API
- Evidence
- Signed pack (Ed25519) + hash-chained record with every output and receipt; /mrm/verify recomputes every number; Markdown binder
- Not
- A validation opinion, legal advice, or a disparate-impact study on real applicants
Pick the tier for the quality you need
Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.
Lite
grader on one 32 GB card (self-host)
The same pinned grader at 4-bit on an RTX 5090; the runner on the same box. Test your own endpoints on your network.
- Models
- decosa-api model-risk module (decosa_api/verticals/mrm)
- Qwen3.8-27B (NVFP4)
- Hardware
- 1x RTX 5090 32 GB (estimate)
- Quality evidence
- Detection and false alarms on this hardwarenot measured yetnot measured yet
- Latency
- estimate: not run on this card
- Verification
- Proof: partialSelf-host onlySelf-hosted calls get receipts signed by your own box (attested), not by the gateway.
- In the hosted demo
Standard
hosted grader and demo systems (Qwen3.8-27B)
What the hosted API runs: every probe of the hosted system and every grader call is receipted by the gateway.
- Models
- decosa-api model-risk module (decosa_api/verticals/mrm)
- Qwen3.8-27B (NVFP4)
- Hardware
- 1x RTX PRO 6000 96 GB (measured, shared)
- Quality evidence
- False alarms on unchanged systems (held-out)0 of 8 ALERT, 0 of 8 WATCH (95% CI for the alert rate 0-32%)docs/evals/model-risk-pack.md
- Injected prompt, bias and model changes caught (held-out)8 of 8 ALERT: vendor prompt update 2/2, ZIP-code bias 2/2, age bias 2/2, swap to Qwen3-1.7B 2/2; each named the right lanedocs/evals/model-risk-pack.md
- Sampling drift caught (temperature 0.3 and 0.7)0 of 4 ALERT (1 WATCH): the triage decisions did not change, so the pack did not alarmdocs/evals/model-risk-pack.md
- Adverse-action drafts lane5/5 in every configuration: it did not separate these injections (weak evidence as built)docs/evals/model-risk-pack.md
- Latency
- measured on a shared GPU through the gateway: one to two minutes for an unchanged pack; longer with a re-check.
- Verification
- Proof: strongGateway-signed receipt per call, embedded in the signed record.
Every model in the stack
| Model | Tiers | Params · VRAM | Verification | Details |
|---|---|---|---|---|
Pack runner: suite, calls, parsing, grading rules, stability, fairness, re-check, model card, signed pack and record (no model; CPU)decosa-api model-risk module (decosa_api/verticals/mrm) 0 GBProof: partial | LiteStandard | 0 GB | Proof: partial | |
| ||||
Fixed grader (typed judgments) and the hosted system under testQwen3.8-27B (NVFP4)nvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab) 27.8B · 20 GBProof: strongIn the hosted demo | LiteStandard | 27.8B · 20 GB | Proof: strongIn the hosted demo | |
| ||||
Injected problem in the eval: the 'vendor' silently moved to a small modelQwen3-1.7B (BF16, CPU)Qwen/Qwen3-1.7B on Hugging Face (opens in a new tab) 1.7B · 0 GBNo proof yetSelf-host only | n/a | 1.7B · 0 GB | No proof yetSelf-host only | |
| ||||
Tools, services and hardware
Tools
- Endpoint auditor golden prompts (vertical 09)Apache-2.0
The identity lane: 10 fixed prompts at temperature 0, compared with the validated run and, for documented open weights, with the auditor's signed reference.
- Typed judgments (vertical 24)Apache-2.0
Grades each adverse-action draft: one yes/no per listed reason, one for an added reason or a protected characteristic.
- Suite fernhill-v1 (synthetic)Apache-2.0
15 complaint-triage cases, 6 counterfactual bases x 5 variants (three names, a ZIP code, an age), 5 adverse-action cases. Written for this demo; the credit union is fictional.
- scripts/mrm_eval.py, docs/evals/model-risk-pack.mdApache-2.0
Injects problems (prompt change, bias, drift, model swap) and runs unchanged systems, on separate dev and test splits.
- scripts/mrm_monitor.pyApache-2.0
Scheduled monitoring from cron or a systemd timer: keeps the baseline and each pack, prints the trend, exits 2 on an alert.
Services
- decosa-api:8445
${DECOSA_REGISTRY}/decosa-api:<tag>GET /mrm/info, /mrm/targets, /mrm/suites/{id}; POST /mrm/runs (SSE or JSON), /mrm/verify, /mrm/render, /mrm/trend. Stores demo packs only.
- vLLM (grader and hosted system):8114
vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1Qwen3.8-27B NVFP4 behind our gateway (hosted) or called directly (self-host).
Hardware
- 1x RTX 5090 32 GB Fits
Qwen3.8-27B NVFP4 needs about 20 GB of weights plus KV cache. Estimate: not run for this vertical.
- 1x RTX PRO 6000 Blackwell 96 GB Fits
Measured on our server: the hosted demo and the eval ran on this card, shared with other services.
- CPU only Fits
The runner needs no GPU. With the grader on another box (or the hosted API), a CPU box can run packs against any endpoint.
Latency per lane
- whole pack, unchanged system (82 calls: 10 identity, 15 triage, 36 fairness, 5 drafts, 16 grader), 2-4 in flight, hosted gateway route62.0 s
Measuredmeasured on our server 2026-09-25: 54-118 s over 8 held-out runs while the GPU was shared with other evaluation jobs
- whole pack with a re-check of the flagged lanes (118-154 calls)135.0 s
Measuredmeasured on our server 2026-09-25: 77-175 s, shared GPU
- whole pack, self-hosted direct route (same GPU, quieter)13.0 s
Measuredmeasured in the self-host sandbox 2026-09-25: 12-14 s for 82 calls
Notes
- Held-out eval: on 8 unchanged runs the pack never alarmed; on 8 runs with a prompt update, a ZIP-code or age bias, or a swap to a small model it said ALERT every time and named the lane and the attribute.
- It did not catch sampling turned up to temperature 0.3-0.7: the triage decisions stayed the same. For a black-box vendor the pack sees behaviour, not settings.
- Policy was set on a separate dev split with different injections; the only change from dev was the identity threshold, and the test set was run once afterwards.
- The stand-in vendor is the same open weights behind a hidden prompt; the swapped-model demo run is a real recorded run against Qwen3-1.7B on CPU, which the hosted box does not keep running.
Run this exact stack on your machine
Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.
# Assemble the Decosa model-risk evidence pack on this machine
You are setting up an evidence-pack runner for the LLMs we use in regulated work (complaint triage, adverse-action
reasons and similar). It re-runs a fixed, signed test suite against a system we use: identity prompts, graded cases,
stability against the validated run, counterfactual fairness pairs. It re-checks anything it flags, and writes a signed
pack plus a hash-chained record that our validators can recompute. Work step by step, show me each command before you
run anything with `sudo`, and stop to ask if a check fails.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/model-risk-pack.zip (3 KB, 10 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py model-risk-pack` (the api image carries the same bundle under /app/rehearsal/model-risk-pack/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py model-risk-pack --bundle model-risk-pack.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the run finishes without errors", "all 14 calls were made (10 identity probes and 4 cases)", "no call failed"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## 0. Ground rules and licences
- Grader and default system: Qwen3.8-27B (Apache-2.0). The runner is decosa-api (AGPL-3.0-or-later) and needs no GPU. Any
system we test is ours or our vendor's; testing it must be allowed by our contract with that vendor.
- Our test cases stay on this machine. Bind every port to 127.0.0.1. Packs for our own endpoints are not stored by the
service: they go back to the caller. Logs carry counts only; do not add request logging.
- The pack is evidence, not a validation opinion, and not legal advice. It detects only what its suite probes.
## 1. Check the machine
1. `nvidia-smi`: one GPU with at least 32 GB for the grader (Qwen3.8-27B NVFP4 needs about 20 GB of weights plus KV
cache; an RTX PRO 6000 96 GB or an RTX 5090 32 GB both work). Driver 570 or newer. On a card without NVFP4, use the
FP8 weights. If we already run an OpenAI-compatible Qwen3.8-27B server, skip the `llm` service and point the api at it.
2. `docker --version` and `docker compose version`. If Docker or the NVIDIA container toolkit is missing, install them
from the official Docker and NVIDIA repositories after asking me, then run
`docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
## 2. Images and weights
- `${DECOSA_REGISTRY}/decosa-api:<tag>` (**publishing soon**). If the pull fails, build from source:
`git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host>` (access required), check out a release that contains
`decosa_api/verticals/mrm/`, and run `docker build -f docker/api/Dockerfile -t decosa-api:local .`
- `vllm/vllm-openai:v0.29.0` with `nvidia/Qwen3.8-27B-NVFP4` (or `Qwen/Qwen3.8-27B-FP8`). Keep the pinned image and
weights: the grader's answers are part of the evidence, so it must not change between runs.
## 3. docker-compose.yml
Write this in `~/decosa/mrm/` (use the `decosa-api:local` image name if you built from source):
```yaml
name: decosa-mrm
services:
llm:
image: vllm/vllm-openai:v0.29.0
command: ["--model", "nvidia/Qwen3.8-27B-NVFP4", "--served-model-name", "qwen3.8-27b", "--max-model-len", "32768",
"--enable-prefix-caching", "--seed", "0"]
ports: ["127.0.0.1:8114:8000"]
volumes: ["~/.cache/huggingface:/root/.cache/huggingface"]
deploy: { resources: { reservations: { devices: [{ driver: nvidia, count: 1, capabilities: [gpu] }] } } }
healthcheck: { test: ["CMD", "curl", "-fs", "http://localhost:8000/v1/models"], interval: 30s, retries: 20 }
api:
image: ${DECOSA_REGISTRY}/decosa-api:<tag>
ports: ["127.0.0.1:8445:8445"]
environment:
DECOSA_HOST: 0.0.0.0
DECOSA_PORT: "8445"
DECOSA_DATA_DIR: /data
DECOSA_LLM_ROUTE: direct
DECOSA_LLM_URL: http://llm:8000/v1
DECOSA_LLM_MODEL: qwen3.8-27b
DECOSA_ADMIN_SECRET: ${DECOSA_ADMIN_SECRET} # for minting keys on this box; keep it in ~/decosa/mrm/.env (0600)
DECOSA_MRM_ALLOW_PRIVATE: "1" # self-host: we may test endpoints on our own network, http included
DECOSA_MRM_MAX_CONCURRENT: "2"
DECOSA_MRM_WORKERS: "4"
volumes: ["mrm-data:/data"] # a named volume: the image runs as a non-root user, so a root-owned bind mount fails
depends_on: { llm: { condition: service_healthy } }
healthcheck: { test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8445/mrm/info', timeout=4)"], interval: 30s, retries: 10 }
volumes:
mrm-data: {}
```
On the first start the api creates this box's Ed25519 key in `/data/attest/` inside the `mrm-data` volume (mode 0600).
Back it up (`docker compose cp api:/data/attest ./attest-backup`, then store that folder safely); never print it. Every model call on the direct route gets a receipt signed with that key (status `attested`: an attestation
by us, the operator, not a proof of computation), and so does every pack. Validators pin the public key from
`GET /attest/signing-key`.
Demo sessions have 20,000 generated tokens, enough for about three packs. For scheduled runs mint a key:
`curl -s -XPOST localhost:8445/v1/keys -H "authorization: Bearer $DECOSA_ADMIN_SECRET" -H 'content-type: application/json' -d '{"label": "mrm", "verticals": ["model-risk-pack"], "daily_llm_tokens": 500000}'`
and keep the `dk_…` key in an environment variable.
## 4. Smoke test
1. `curl -s localhost:8445/mrm/info | jq '{policy, grader: .grader.prompt_sha256, suites}'` and
`curl -s localhost:8445/mrm/targets | jq '.[] | {id, available, baseline}'`: `lender` and `vendor` must be available.
2. Token: `T=$(curl -s -XPOST localhost:8445/demo/session -H 'content-type: application/json' -d '{"vertical":"model-risk-pack"}' | jq -r .token)`.
3. `curl -s -XPOST localhost:8445/mrm/runs -H "authorization: Bearer $T" -H 'content-type: application/json' -d '{"target":"lender"}' > run.json`
(about 80 calls; 1-5 minutes). Expect `.pack.summary.performance.classify.correct.k` of 14-15 of 15, drafts 4-5 of 5,
0-1 fairness flips, and `.pack.verdict.status` `pass`. The shipped baseline was recorded on the hosted stack, so a
self-hosted engine can show some identity or stability difference: that is the check working. Record your own
baseline (step 5) before trusting the verdicts.
4. `jq '{pack, record}' run.json | curl -s -XPOST localhost:8445/mrm/verify -H 'content-type: application/json' -d @-`
must give `ok: true`, `record.ok: true` and `recompute.ok: true`. Change one `text` in a record entry and verify
again: it must fail at that entry.
5. Run it on our own system: an OpenAI-compatible endpoint (`{"endpoint": {"base_url": "http://host:port/v1", "model": "…", "black_box": true}}`
when the endpoint applies its own prompt), with our own suite in the same schema (`GET /mrm/suites/fernhill-v1` is
the template: classify cases with expected values, fairness bases with {slots}, drafts with their listed reasons).
Keep the first pack as the baseline and send it as `baseline` on every later run.
## 5. Schedule it
Copy `scripts/mrm_monitor.py` from the repo (standard library only) and run it weekly from cron or a systemd timer:
`DECOSA_API_KEY=dk_… python3 mrm_monitor.py --api http://127.0.0.1:8445 --dir ~/decosa/packs/vendor --target vendor`.
It keeps the first pack as `baseline.json`, stores each pack, record and Markdown binder, prints the trend, and exits
2 on an alert, so the scheduler can page someone. Point the site at this box with
`NEXT_PUBLIC_DECOSA_API=http://127.0.0.1:8445` to use the console. Contract: `API_CONTRACT.md`, "Model-risk evidence pack".
Off by default. Never on a box that holds confidential test cases. Only if I ask, and only after my explicit yes,
follow the provider guide at `/provide` on the site.Rules and regulations it checks againstDated, linked to the primary source; not legal advice
Regulation watch
Loading the watch status…
3 laws, rules and guidance pages cited; 3 watched nightly at the primary source. A change marks this page for a human re-check; nothing is edited automatically. What we cite and how it is watched
Technical detailsModels, where it runs, labels
In short
Last reviewed
- What it is
- Signed evidence that an LLM your bank, insurer or lending team relies on still behaves as validated, with a record your validators can recompute.
- Who it's for
- Teams in finance and insurance and compliance and trust.
- Where it runs
- Hosted or self-host
- Key numbers
- 0 of 8 / 0 of 8 False alarms on unchanged runs (ALERT / WATCH) (test split, n = 8)
- 8 of 12 runs Injected changes detected (ALERT) (test split, n = 12)
- 0 of 4 runs Sampling drift detected (temperature 0.3 / 0.7) (test split, n = 4)
- 62.0 s Median end-to-end run, hosted (QA sweep 2026-09-25)
- Models
- Qwen3.8-27B (hosted system and fixed grader); any OpenAI-compatible endpoint under test
- Where
- Hosted or self-host
- Checks
- Receipt per call; signed pack and hash-chained record that recompute
- Industry
- Finance and insurance · Compliance and trust
- Input
- Endpoint or URL
- Output
- Signed record or verdict · Notes, reports and drafts
- Data
- Confidential business data
- Hardware
- 1× 96 GB GPU
- Licence
- Permissive (Apache-2.0, MIT)
- Runs in
- Decosa hosted · Self-host
- Built from
- Endpoint audit · Typed judgment · Signed record
Questions people ask
Does SR 26-2 cover LLMs and generative AI?
Not directly. SR 26-2 (Fed, OCC, FDIC, 17 Apr 2026) replaced SR 11-7 and OCC 2011-12, and its footnote 3 puts generative and agentic AI models outside its scope, leaving their controls to each bank's own risk management and governance. The Decosa Model-risk evidence pack is evidence for that governance; it is not a validation opinion and does not by itself meet the rule. Not legal advice.
What does the Model-risk evidence pack test?
The Model-risk evidence pack re-runs a fixed suite against an LLM endpoint: your own open-model deployment, a vendor's black-box API, or any OpenAI-compatible endpoint. It has four lanes: identity probes against the validated run, graded canary cases, stability on the same inputs, and counterfactual fairness pairs that change only a name, ZIP code or age. A flagged lane is re-run independently before the pack says alert.
How well does the Model-risk evidence pack detect a changed vendor model?
On the held-out test, the Model-risk evidence pack alerted on 8 of 8 prompt, bias and model changes, each naming the right lane, and raised no alert on 8 unchanged runs, a 0-32% interval on the false-alarm rate. It detected 0 of 4 sampling-temperature changes, and the drafts grader did not separate any injection. The suite and injections are synthetic and were written by the team that built the pack.
Can validators re-check the Model-risk evidence pack's numbers?
Yes. Each Model-risk evidence pack is signed (Ed25519) and comes with a hash-chained record of every output and its receipt, and /mrm/verify recomputes every number from that record; in the self-host check, a one-word edit in the record failed verification at that entry. The pack also includes a model card, a trend line and a Markdown binder, at about $0.013 per pack (82 calls) at the gateway list price.
Does the fairness lane measure disparate impact on applicants?
No. The Model-risk evidence pack's fairness probes change one attribute at a time among the values in the suite, so a bias on an unprobed attribute or value would pass, and they do not measure disparate impact on real applicants. Lenders still need specific principal reasons for adverse action under Regulation B (12 CFR 1002.9); the drafts grader checks that listed reasons appear, not how specific they are.
Ask a question or leave feedbackWe read every message and publish useful answers
Ask about Model-risk evidence pack
We read every message. Questions, comments and our answers show here once we have reviewed and approved them.
Loading questions…