Skip to content
decosa
LiveHostedSelf-hostMac

Catch AI answers your sources don't back

Every sentence of an AI-drafted answer checked against the source passage it relies on, with the quote, and a pass, flag or block you can put in front of Send.

Held-out test0.966Unsupported sentences caught, strict gate (it flags about 30% of sentences for a look; held-out)
On production17 smedian on production (2026-09-25); slower when the service is busy
List price~$0.20 per 100 checksmeasured, at list price

Built on: Grounding, Signed record

Loading the tool…

Use it your way

Use it from your codeThe hosted API with your key, and prompts to paste into a coding agent
Hosted · by Decosa

Get an API key

  • Call the grounding check API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Nothing to install; we run the models.
Self-host · your GPUs

Run it yourself, on request

  • The same open models and app, on 1× RTX PRO 6000 (96 GB) or 1× RTX 5090 (32 GB) for the judge; the checker itself runs on CPU.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.

Build with it

Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.

Base URL
https://api.decosa.ai
Auth
Authorization: Bearer $DECOSA_API_KEY (or a demo session token)
Tool id
grounding

Use the hosted API

# Decosa Grounding check: use the hosted API

You are wiring Decosa's grounding check into this project. It takes a text (an answer, an email, a FAQ, a summary) and
the sources it is supposed to rest on, and says for each sentence whether the sources support it: `supported`,
`partial`, `unsupported` or `contradicted`, with the source span it rests on and a confidence. A claims gate turns that
into pass, flag or block before the text ships. Every verdict is one model call with its own signed receipt, and the
check ends with a report signed by the server. Use only what is listed below. If you need something else, stop and ask me.

- Base URL: `https://api.decosa.ai`
- Health check: `GET https://api.decosa.ai/healthz`.
- "Supported by the sources" is not "true": the check reads only the sources you send. Say so wherever you show results.

## Auth: API key (or a demo session)
1. Preferred: an API key (`dk_…`) from "Get an API key" on the tool page. Keep it in an environment variable,
   `DECOSA_API_KEY`, never in code. Send `Authorization: Bearer $DECOSA_API_KEY`.
2. Without a key: `POST https://api.decosa.ai/demo/session` with `{"vertical": "grounding"}` returns `{"token", "expires_at", "budget"}`.
   a limited number of sessions per network per hour (the current limits are in `demo_sessions` of GET /healthz); each session has a token allowance (its `budget`) (about 120 sentences). Over a limit: HTTP 429 with `Retry-After`.
3. A check needs about 160 generated tokens of budget per sentence (402 otherwise). One check at a time per demo token.

## Endpoints
- `POST /grounding/check` (token). Body:
  `{"text": "...", "sources": [{"title": "...", "text": "..."}], "question"?: "...", "gate"?: {"block_on": ["contradicted", "unsupported"], "flag_on": ["partial"], "min_confidence": 0}}`
  - Limits: text up to 8,000 characters and 40 sentences; up to 10 sources and 60,000 characters in total. Send text, not URLs:
    fetch pages yourself.
  - JSON by default: `{decision, counts, blocking, flagged, clean_text, sentences: [{i, text, verdict, certainty, confidence, evidence: [{span, source, start, end, role, quote}], reason, receipt_ids}], receipts, report, budget, note}`.
  - With `Accept: text/event-stream` (or `"stream": true`) it streams: `ready` (sentence and span offsets), then a `receipt`
    and a `sentence` event per sentence as each finishes (not in text order), then `report`, `budget`, `done`.
- `POST /grounding/gate` (token): same body, always JSON → `{decision: "pass"|"flag"|"block", issues: [{i, text, verdict, confidence, reason, evidence, action}], clean_text, counts, report, receipts}`.
  The default policy blocks `contradicted` and `unsupported`, flags `partial`, and blocks any sentence that could not be judged.
- `POST /grounding/verify` (no token) `{"report": {...}, "text"?: "...", "sources"?: [...]}` → `{valid_signature, signed_by_this_server, text_matches?, sources_match?}`.
- `GET /grounding/info`, `GET /grounding/samples`, `GET /attest/signing-key` (no token).

Offsets (`start`, `end`) are Python string indices (Unicode code points) into the text or source exactly as sent. In
JavaScript, index with `Array.from(text)` if the text can contain emoji or other characters outside the BMP.

## Example: block an outbound email that says things the profile does not (Python, `pip install httpx`)
```python
import httpx, os
API = "https://api.decosa.ai"
H = {"Authorization": f"Bearer {os.environ['DECOSA_API_KEY']}"}
body = {"text": email_draft, "sources": [{"title": "Candidate profile", "text": profile}, {"title": "Role brief", "text": role}]}
r = httpx.post(f"{API}/grounding/gate", json=body, headers=H, timeout=180)
r.raise_for_status()
g = r.json()
if g["decision"] == "block":
    for issue in g["issues"]:
        print(issue["action"], issue["verdict"], issue["confidence"], issue["text"], "--", issue["reason"])
    raise SystemExit("not sending: fix the flagged sentences")
```

## Verify a report yourself (`pip install cryptography`)
```python
import hashlib, json, urllib.request
from cryptography.hazmat.primitives.asymmetric.ed25519 import Ed25519PublicKey
pub = json.load(urllib.request.urlopen("https://api.decosa.ai/attest/signing-key"))["pubkey"]
assert report["signer"] == pub
body = {k: v for k, v in report.items() if k != "sig"}
msg = report["v"].encode() + b"\n" + json.dumps(body, sort_keys=True, separators=(",", ":"), ensure_ascii=False).encode()
Ed25519PublicKey.from_public_bytes(bytes.fromhex(pub)).verify(bytes.fromhex(report["sig"]), msg)
assert hashlib.sha256(email_draft.encode()).hexdigest() == report["text"]["sha256"]
```
The report carries hashes and offsets, never the text, so you can store it next to the sent email as evidence of the check.
Each id in `report["receipts"]` resolves at `GET https://api.decosa.ai/receipts/{id}` to the signed model receipt for that verdict.

## Honest limits
- On the RAGTruth benchmark the judge finds most sentences that human annotators marked as unsupported, and also flags
  sentences they left alone: see the numbers on the Stack tab before you let the gate block without a human look.
- Confidence is the measured agreement of that verdict and certainty with human labels on a dev set, not a guarantee.

Run it yourself (containers)

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

# Decosa Grounding check: run it yourself (containers)

You are setting up the Decosa grounding check on this machine, so unpublished drafts and their sources never leave it.
It checks each sentence of a text against the sources I give it and signs a report. Nothing is sent to Decosa's hosted API.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/grounding.zip (2 KB, 10 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py grounding` (the api image carries the same bundle under /app/rehearsal/grounding/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py grounding --bundle grounding.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the claims gate blocks the answer", "every claim sentence got a verdict (none errored)", "the battery-life sentence is contradicted or not backed"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
   instructions for this distribution (docs.docker.com/engine/install). For the GPU judge, also install the NVIDIA
   container toolkit and check `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
   `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
   Read it. Keep the `llm` service (Qwen3.8-27B on vLLM) and the `api` service. For the `api` service set
   `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b` and bind every port to
   127.0.0.1. Without a GPU, set `DECOSA_GROUNDING_NLI=cross-encoder/nli-deberta-v3-base` instead and use `"judge": "nli"`
   in requests (CPU only, lower accuracy, no model receipts).
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check (the first start
   downloads about 20 GB of weights).
4. Check: `curl -fsS http://127.0.0.1:<PORT>/grounding/info` lists the verdicts and the judge; `GET /attest/signing-key`
   shows this box's public key. Show me the key: it is what others pin to verify my reports.
5. Smoke test: get a token with `POST /demo/session {"vertical":"grounding"}`, fetch `GET /grounding/samples`, and send the
   `support-answer` sample to `POST /grounding/gate`. Expect `decision: "block"` with the battery-life and 240 V sentences
   as `contradicted`, and the reset steps as `supported`. Then `POST /grounding/verify` with the report: `valid_signature`
   and `signed_by_this_server` must both be true.
6. Report back: the public key and key id, the smoke-test decision, the verdicts and how long the check took.

Off by default. Joining as a provider serves other people's requests on this GPU; never do it on a box that holds
confidential drafts. If I ask for it later, follow the Provide page instead of improvising.

No NVIDIA GPU? This tool also runs entirely on an Apple Silicon Mac (MLX, 32 GB of unified memory or more): use https://decosa.ai/prompts/grounding-mac.md instead.
Run it on your own hardwareWhat it needs, and the prompt that sets it up

Run it on your own GPU

Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.

  • CPU only, 64 GB RAMlite tierRuns with a smaller tier

    The standard tier does not fit: Qwen3.8-27B (NVFP4) needs a GPU. The lite tier fits.

  • GeForce RTX 4090standard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)

  • GeForce RTX 5090standard tierRuns

    The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

  • 2x GeForce RTX 5090standard tierRuns

    The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size).

  • L40Sstandard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.

  • H100 80 GB (SXM)standard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.

  • RTX PRO 6000 Blackwell 96 GBstandard tierRuns

    The standard tier fits (57.6 of 96 GB).

  • 2x RTX PRO 6000 Blackwell 96 GBstandard tierRuns

    The standard tier fits (57.6 of 192 GB).

  • Apple M3 Ultra (Mac Studio), 96 GBstandard tierRuns

    The standard tier fits (32 of 96 GB).

  • Apple M5 Max, 64 GBstandard tierRuns

    The standard tier fits (32 of 64 GB).

Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

  1. 1

    Check the GPU, Docker and the NVIDIA Container Toolkit

    The driver must see the GPU, and Docker must be able to pass it into a container.

    nvidia-smi
    docker compose version
    docker run --rm --gpus all ubuntu nvidia-smi
  2. 2

    Fetch the compose file

    One file describes the API and the language model as services.

    mkdir -p ~/decosa && cd ~/decosa
    curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml
  3. 3

    Pull and start

    The first start downloads pinned model weights, tens of gigabytes.

    docker compose pull
    docker compose up -d
  4. 4

    Check health

    Wait until the API reports ok with the language model loaded. Then point your app at the local base URL.

    curl -fsS http://localhost:<PORT>/healthz
    # {"ok": true, "llm": true, ...}
    curl -fsS -X POST http://localhost:<PORT>/demo/session \
      -H 'Content-Type: application/json' -d '{"vertical":"grounding"}'

Set up with a coding agent, rehearse on mock data, then go private

  1. Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
  2. Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's expected.json. Every check must print PASS.
  3. Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
Rehearsal command
docker compose exec api python scripts/rehearse.py grounding

Download the mock-data bundle (2 KB, 10 checks)expected.json

A support answer about a fictional thermostat, checked sentence by sentence against its manual. Two claims conflict with the manual (battery life, 240 V heaters), so the gate must block it, and the signed report must verify and catch a changed verdict.

What the rehearsal checks
  • the claims gate blocks the answer
  • every claim sentence got a verdict (none errored)
  • the battery-life sentence is contradicted or not backed
  • the 240 V heater sentence is contradicted or not backed
  • the factory-reset sentence is supported
  • the signed report verifies against the same text and sources
  • the verified report matches the text
  • the verified report matches the sources
  • a report with one verdict changed no longer verifies
  • every model call has a signed receipt

Licence: Synthetic: a fictional product (Halden T3) and manual written for Decosa, no real people or products. Part of decosa-api, AGPL-3.0-or-later.

Prompt for your coding agent

# Decosa Grounding check: run it yourself (containers)

You are setting up the Decosa grounding check on this machine, so unpublished drafts and their sources never leave it.
It checks each sentence of a text against the sources I give it and signs a report. Nothing is sent to Decosa's hosted API.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/grounding.zip (2 KB, 10 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py grounding` (the api image carries the same bundle under /app/rehearsal/grounding/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py grounding --bundle grounding.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the claims gate blocks the answer", "every claim sentence got a verdict (none errored)", "the battery-life sentence is contradicted or not backed"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
   instructions for this distribution (docs.docker.com/engine/install). For the GPU judge, also install the NVIDIA
   container toolkit and check `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
   `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
   Read it. Keep the `llm` service (Qwen3.8-27B on vLLM) and the `api` service. For the `api` service set
   `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b` and bind every port to
   127.0.0.1. Without a GPU, set `DECOSA_GROUNDING_NLI=cross-encoder/nli-deberta-v3-base` instead and use `"judge": "nli"`
   in requests (CPU only, lower accuracy, no model receipts).
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check (the first start
   downloads about 20 GB of weights).
4. Check: `curl -fsS http://127.0.0.1:<PORT>/grounding/info` lists the verdicts and the judge; `GET /attest/signing-key`
   shows this box's public key. Show me the key: it is what others pin to verify my reports.
5. Smoke test: get a token with `POST /demo/session {"vertical":"grounding"}`, fetch `GET /grounding/samples`, and send the
   `support-answer` sample to `POST /grounding/gate`. Expect `decision: "block"` with the battery-life and 240 V sentences
   as `contradicted`, and the reset steps as `supported`. Then `POST /grounding/verify` with the report: `valid_signature`
   and `signed_by_this_server` must both be true.
6. Report back: the public key and key id, the smoke-test decision, the verdicts and how long the check took.

Off by default. Joining as a provider serves other people's requests on this GPU; never do it on a box that holds
confidential drafts. If I ask for it later, follow the Provide page instead of improvising.

No NVIDIA GPU? This tool also runs entirely on an Apple Silicon Mac (MLX, 32 GB of unified memory or more): use https://decosa.ai/prompts/grounding-mac.md instead.

Help me customise for my hardware

Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.

Hardware

GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page

RunsGrounding check on GeForce RTX 5090: use the Standard · one GPU for the judge (hosted demo) tier

The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

What this tool's stack says about this hardware:

  • 1x RTX 5090 32 GB (fits): Qwen3.8-27B NVFP4 needs about 20 GB of weights plus KV cache. Estimate: same stack as the code use case, not run here for this vertical.

Standard · one GPU for the judge (hosted demo): what changesuses estimates

  • Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
Memory per component
  • Checker: decosa-api grounding module (decosa_api/verticals/grounding). CPU. Runs on CPU (vram_gb 0 in stack.json).
  • Judge: Qwen3.8-27B (NVFP4). ~57.6 GB (at least ~28 GB), weights 21.4 GB (from stack.json). Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)

Expected speed

Not measured.

Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.

Setup prompt for this hardware

The self-host prompt for Grounding check, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.

# Set up Grounding check on my hardware

Fetch https://decosa.ai/prompts/grounding-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied.

## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=grounding)

Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4).
Quality tier: Standard · one GPU for the judge (hosted demo) (standard). Fit check: runs with changes, about 28 GB of 32 GB used; some memory numbers are estimates, not measurements.

First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything.

Use these components (the setup below describes the standard tier; change it to match):
- Checker: decosa-api grounding module (decosa_api/verticals/grounding), CPU
- Judge: Qwen3.8-27B (NVFP4) (nvidia/Qwen3.8-27B-NVFP4), 57.6 GB. Change: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

GPU placement (set each service's device and its vLLM --gpu-memory-utilization to about the share shown):
- GPU 0: Qwen3.8-27B (NVFP4) ~28 GB (88%); about 4 GB left

During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed.

The stack's own component list and compose layout: https://decosa.ai/prompts/grounding-assemble.md

Or on a Mac Studio

No NVIDIA GPU needed: every model this tool uses runs natively on Apple Silicon through MLX. Any M-series Mac with 32 GB of unified memory or more. Measured speeds and what runs where

From a checkout of decosa-api, one command sets up the models and the API: scripts/mac/setup.sh

Mac prompt for your coding agent

# Decosa Grounding check: run it on this Mac (Apple Silicon, no NVIDIA GPU)

You are setting up the Decosa Grounding check on this Mac, natively on Apple Silicon. The models run on the Mac's GPU
through MLX and decosa-api runs from a git checkout with `uv`. Docker is not used for the models, because Docker on
macOS cannot reach the GPU. Nothing is sent to Decosa's hosted API.

Every model this tool needs runs on the Mac. It needs 32 GB of unified memory or more.

Ask me before any command that needs sudo or installs software with Homebrew, and show me the command first. Never stop
or kill a process this setup did not start; if a port is taken, pick another one.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/grounding.zip (2 KB, 10 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `.venv/bin/python scripts/rehearse.py grounding` in the decosa-api checkout (the key comes from ~/.decosa-mac/api.key).
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the claims gate blocks the answer", "every claim sentence got a verdict (none errored)", "the battery-life sentence is contradicted or not backed"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## What runs where

| Part | On an NVIDIA GPU | On this Mac | Status |
|---|---|---|---|
| Checker: segmentation, evidence selection, gate, signed report (no model; CPU) | Python on CPU | The same Python module, run with uv | Runs, measured |
| Judge: one call per sentence, verdict and cited spans | NVFP4 on vLLM 0.29 (Blackwell) | MLX 4-bit (EigenLabs/Qwen3.8-27B-4bit) on mlx_lm.server 0.31.3; oMLX 0.6.1 with MTP as an option | Runs, measured |

## Steps
1. Check the machine: `uname -m` must print `arm64` (an M-series chip; Intel Macs cannot run MLX), and
   `sysctl -n hw.memsize` should be at least 32 GB for this tool. Check about 30 GB of free disk with
   `df -h ~`. Show me the chip (`sysctl -n machdep.cpu.brand_string`) and the memory.
2. Tools: `uv --version`. If it is missing, ask me, then `brew install uv`.
3. Code: `git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host> ~/decosa-api` (access required) and `cd ~/decosa-api`.
   Check that `scripts/mac/setup.sh` exists; if it does not, the checkout is too old: stop and tell me.
4. Start everything with one command: `scripts/mac/setup.sh`. It creates `.venv` (decosa-api)
   and `.venv-mac` (MLX, mlx-lm, mlx-audio), downloads the weights with the Hugging Face CLI (about 16 GB for the
   language model), starts the model servers and decosa-api on 127.0.0.1, and mints a local API key
   into `~/.decosa-mac/api.key` (mode 0600). The first run takes a while because of the downloads; later runs reuse them.
   If a download fails with 401 or 403, ask me for a Hugging Face token and set `HF_TOKEN`.
5. Check health: `scripts/mac/setup.sh status` shows each server, and `curl -fsS http://127.0.0.1:8445/healthz` must
   report `"llm": true`. `curl -fsS http://127.0.0.1:8445/attest/signing-key` shows this Mac's public key:
   show it to me, because it is what others pin to check the receipts and records this Mac signs.
6. Smoke test: `.venv/bin/python scripts/mac/bench_usecases.py grounding`. It runs the tool's own sample end to end
   against the local API with the local key and prints `ok`, the wall time, the model calls and the receipts.
   `ok=True` is the pass condition. If it fails, read `~/.decosa-mac/logs/*.log` and tell me what you found.
7. Point the app at it: the API is `http://127.0.0.1:8445` with `Authorization: Bearer $(cat ~/.decosa-mac/api.key)`,
   the same routes as the hosted API. To stop everything: `scripts/mac/setup.sh stop`.
8. Report back: the chip and memory, the public key, the smoke-test result and its time, and the output of
   `scripts/mac/setup.sh status`.

## Good to know
- Receipts: every model call is signed with this Mac's own Ed25519 key and names the exact MLX weights
  (`qwen3.8-27b-mlx-4bit` with a hash of the downloaded files). There is no gateway countersignature on a
  self-hosted Mac.
- The weights are a 4-bit MLX build of the same open models, not the NVFP4 build the hosted route and the published
  evals use. Expect small differences in wording and scores.
- Faster drafting: `scripts/mac/setup.sh stop && scripts/mac/setup.sh --engine omlx` serves the
  model with oMLX and multi-token prediction (about 2x faster for a single long answer, no faster for many parallel
  calls; typed judgments then use sampling because oMLX returns no log-probabilities).
- Measured speeds for a Mac Studio M3 Ultra and the memory each tool needs: https://decosa.ai/mac. Full details:
  `docs/self-host-mac.md` in the checkout.

The proof

How we tested itEval results and end-to-end checks, hosted and self-hosted, with dates

Verified end to end

Hosted: verified 25 Sep 2026 · measured 25 Sep 2026: · p50 17 s · ~$0.002 per run · 5 receipts

Loading the nightly status…

Self-host: verified 25 Sep 2026 · Fresh git clone of decosa-api, image built from docker/api/Dockerfile, compose up on 127.0.0.1, sample run end to end against local model servers

Measured cost to run: about $0.20 per 100 checks (hosted, 25 Sep 2026). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.

Verified on 2026-09-25, option A: the api starts, the thermostat sample blocks with the battery and 240 V sentences contradicted and the reset sentence supported at S1.3, the stream and the signed report behave as documented (p50 1.9 s) against a local Qwen3.8-27B vLLM equivalent to the documented one; model-server startup itself not re-verified. Option B (CPU NLI judge) also ran: it blocked the sample but marked the battery-life sentence supported.

Known limits (5)
  • The judge reads only the sources given: "supported" is not "true".
  • Each sentence is one model call that re-sends the sources, so tokens grow with sentences x source length (about 5.8k for a five-sentence answer against a one-page manual).
  • The CPU NLI judge (self-host option B) is coarser: on the thermostat sample it missed one of the two contradictions.
  • Sources must be pasted text; a bare URL is refused (nothing is fetched).
  • Hosted timings were measured on 2026-09-25 while the gateway was degraded under QA load; the same calls took 1-3 s self-hosted. Token counts on hosted receipts are the gateway's metering, which on that date overstated prompt tokens by about 25-80% against the model's tokenizer (a fix is in progress).

Eval results, nightly checks and cost per runVerify a run

How it's builtThe steps, the models and what each one checks
Hosted · by Decosa

Get an API key

  • Call the grounding check API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Nothing to install; we run the models.
Self-host · your GPUs

Run it yourself, on request

  • The same open models and app, on 1× RTX PRO 6000 (96 GB) or 1× RTX 5090 (32 GB) for the judge; the checker itself runs on CPU.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
The open stack

Is each sentence supported by the sources? Per-sentence verdicts with the source span, a claims gate, and a signed report.

Send a text and the sources it is supposed to rest on. The checker splits the text into sentences with exact offsets, numbers the source spans, and asks an open judge model about each sentence in its own call: supported, partial, unsupported or contradicted, with the spans it rests on and a confidence measured on a benchmark. The claims gate turns that into pass, flag or block before an email, FAQ or answer ships, and every check ends with an Ed25519-signed report of hashes and offsets. It is the clinical scribe's sentence verifier, made general, and a module other tools can import.

Deployment
Hosted or self-host
Regulatory
A grounding check is a quality control, not a fact check: it reads only the sources you send, and a sentence copied from a wrong source passes. Verdicts come from a language model and can be wrong in both directions (see the RAGTruth numbers below). The hosted demo keeps no text: drafts and sources stay in memory for the request, the report holds hashes only, and logs carry counts. For unpublished or confidential drafts, self-host. Where rules require substantiating marketing claims (for example the FTC's advertising substantiation policy in the US), a pass here is a record that a check ran against the sources you chose, not proof that a claim is substantiated. Not legal advice. Model licences: Apache-2.0 (Qwen3.8-27B; cross-encoder/nli-deberta-v3-base). Checked 24 Sep 2026.
Architecture
Text description

A text, its sources and an optional question go to the checker, which splits the text into sentences, numbers the source spans, picks the evidence for each sentence and skips lines with no claim. The Qwen3.8-27B judge makes one call per sentence and returns a verdict and the spans it cites; a CPU NLI cross-encoder can stand in on self-hosted machines without a GPU. The checker applies the gate policy and signs a report of hashes and offsets. Outputs: per-sentence verdicts, a pass, flag or block decision with the text minus blocked sentences, and the signed report. On the hosted route each judge call gets a signed receipt that our gateway countersigns. In self-host mode the checker and the judge run on your machine.

Architecture

At a glance

Data retention
Nothing stored: text and sources live in memory for the request. The signed report carries hashes and character offsets, not your text.
Inputs
Text up to 8,000 characters and 40 sentences; up to 10 sources and 60,000 characters in total, as text (files are read in your browser on the console).
Typical hosted cost
A fraction of a cent for a short answer against a one-page source at the gateway's list price (measured). Each run shows its own measured cost.
What leaves the box (self-host)
Nothing: the judge runs on your vLLM (or on CPU with the NLI option) and reports are signed with this box's key.
Quality tiers

Pick the tier for the quality you need

Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.

  • Lite

    CPU only, no GPU

    The checker plus an NLI cross-encoder. Runs anywhere; far less accurate, and no model receipts.

    Models
    • decosa-api grounding module (decosa_api/verticals/grounding)
    • cross-encoder/nli-deberta-v3-base
    Hardware
    Any Linux or macOS machine with Python 3.11+
    Quality evidence
    • RAGTruth test, sentence level: precision / recall on unsupported0.133 / 0.792 (F1 0.227)docs/evals/grounding.md, 300 held-out responses, threshold from dev
    • RAGTruth test: agreement with human labels53.6% (κ 0.09)docs/evals/grounding.md
    • Lexical-overlap baseline, same test0.177 / 0.455 (F1 0.255), agreement 77.1%docs/evals/grounding.md
    Latency
    measured: a fraction of a second per sentence on a few CPU threads.
    Verification
    No proof yetSelf-host onlyNo model call, so no receipt; the report is still signed by the box.
  • In the hosted demo

    Standard

    one GPU for the judge (hosted demo)

    Qwen3.8-27B judges each sentence in its own receipted call. This is what the hosted API runs.

    Models
    • decosa-api grounding module (decosa_api/verticals/grounding)
    • Qwen3.8-27B (NVFP4)
    Hardware
    1x RTX PRO 6000 96 GB (measured) or 1x RTX 5090 32 GB (estimate)
    Quality evidence
    • RAGTruth test, default gate (block = unsupported or contradicted): precision / recall on unsupported0.593 / 0.556 (F1 0.574)docs/evals/grounding.md, 2,069 sentences, 178 unsupported, held out
    • RAGTruth test, default gate: agreement with human labels92.9% (κ 0.535)docs/evals/grounding.md
    • RAGTruth test, strict (partial also flagged): precision / recall0.276 / 0.966 (F1 0.429), agreement 77.9%docs/evals/grounding.md
    • RAGTruth test, response level, default gate: precision / recall0.706 / 0.675docs/evals/grounding.md
    Latency
    measured: seconds per short check on a quiet gateway; much longer when the shared gateway is busy.
    Verification
    Proof: strongEvery verdict is a separate gateway call with a gateway-signed receipt; the signed report lists them all.
  • Needs more compute

    Wanted: the best setup

    a panel of the largest open judges

    DeepSeek-V4.1-Flash and GLM-5.3-Flash judge each sentence beside Qwen3.8-27B. A sentence counts as supported only when the panel agrees, and disagreements are shown. Not served yet.

    Models
    • decosa-api grounding module (decosa_api/verticals/grounding)
    • Qwen3.8-27B (NVFP4)
    • DeepSeek-V4.1-Flash
    • GLM-5.3-Flash
    Hardware
    Network providers: an 8x H200-class node for DeepSeek-V4.1-Flash (476 GB of weights); 2x 96 GB cards or a Mac with 192 GB or more for GLM-5.3-Flash (about 170 GB). The 27B stays on one card. Estimate.
    Quality evidence
    • sentence-level F1 on the grounding set, same protocol as standardnot measured yet
    Latency
    not measured yet
    Verification
    No proof yetSelf-host onlyNot hosted yet, so no receipts today.
    Not served yet. It needs more than one 96 GB card, so it runs on your own bigger box.

Also runs on

  • Fast option-token judgeGemma-4-26B-A4B-itnot servedA 4B-active judge read through option-token probabilities: cheaper per call, not better. Page 32 suggests the dense Qwen3.5-9B (Apache-2.0) instead. The real blocker is logprob passthrough on the gateway. Hardware: 1x RTX PRO 6000 96 GB (BF16 weights are 49 GB).

We host these ourselves when needed: small models get more of our own compute unless we detect a shortage, so they need no community providers.

Components

Every model in the stack

Models in this stack. Each row has a button that shows its licence, engine, verification and evidence.
ModelDetails
Checker: segmentation, evidence selection, gate, signed report (no model; CPU)decosa-api grounding module (decosa_api/verticals/grounding)
0 GBProof: partial
Judge: one call per sentence, verdict and cited spansQwen3.8-27B (NVFP4)nvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab)
27.8B · 20 GBProof: strongIn the hosted demo
Lite judge: NLI cross-encoder on CPU (self-host only)cross-encoder/nli-deberta-v3-basecross-encoder/nli-deberta-v3-base on Hugging Face (opens in a new tab)
184M · 0 GBNo proof yetSelf-host only
Fast judge with option-token probabilities (alternate)Gemma-4-26B-A4B-itgoogle/gemma-4-26B-A4B-it on Hugging Face (opens in a new tab)
26B (4B active) · 49 GBNo proof yetSelf-host only
Second judge: is each sentence supported by its sources?DeepSeek-V4.1-Flashdeepseek-ai/DeepSeek-V4.1-Flash on Hugging Face (opens in a new tab)
552B backbone (763B incl. Engram tables) (8B in / 16B out active) · about 476 GB (estimate)No proof yetSelf-host only
Third judge, from another familyGLM-5.3-Flashzai-org/GLM-5.3-Flash on Hugging Face (opens in a new tab)
321B (18B active) · about 170 GB (estimate)No proof yetSelf-host only

Around the models

Tools, services and hardware

Tools

  • The eval: human span labels of unsupported and conflicting content in QA, summary and data-to-text responses. We used a 300-response sample of its test split, held out.

  • scripts/grounding_eval.py and docs/evals/grounding.mdApache-2.0

    Rebuilds the dev and test samples, runs the judge and both baselines, and writes the metrics and the confidence table.

  • POST /grounding/verifyApache-2.0

    Checks a report's signature against this server's key and, if you send them, the text and source hashes. The console also checks the signature in your browser with WebCrypto.

Services

  • decosa-api:8445
    ${DECOSA_REGISTRY}/decosa-api:<tag>

    GET /grounding/info, /grounding/samples; POST /grounding/check (SSE or JSON), /grounding/gate, /grounding/verify. Keeps no text.

  • vLLM (judge):8114
    vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1

    Qwen3.8-27B NVFP4 behind our gateway (hosted) or called directly (self-host).

Hardware

  • Any CPU, no GPU Fits

    Lite tier only (NLI cross-encoder): 2,996 sentences took 658 s on 8 CPU threads (about 0.2 s each).

  • 1x RTX 5090 32 GB Fits

    Qwen3.8-27B NVFP4 needs about 20 GB of weights plus KV cache. Estimate: same stack as the code tool, not run here for this vertical.

  • 1x RTX PRO 6000 Blackwell 96 GB Fits

    Measured on our server: the hosted demo and the eval ran on this card, shared with other services.

Latency per lane

  • check of 4-5 judged sentences, hosted gateway route1.4 s

    Measuredmeasured on our server 2026-09-24: 0.9-3.0 s for the four samples with a quiet gateway; 12-29 s earlier the same day while other evals saturated the shared gateway

  • one judge call, 6 in flight, direct route2.0 s

    Measuredmeasured on our server 2026-09-24: 1,995 calls in 679 s during the RAGTruth test run (GPU shared with other work)

  • one sentence, NLI cross-encoder on CPU220 ms

    Measuredmeasured on our server 2026-09-24: 2,996 sentences in 658 s, 8 threads

Notes

  • Two operating points from one model. Counting 'partial' as flagged finds 97% of the sentences RAGTruth annotators marked but flags 30% of all sentences; many of those are added details the annotators left alone. So the default gate blocks only unsupported and contradicted (precision 0.59, recall 0.56 on the held-out test) and flags partial for a human look.
  • Confidence is per verdict, from the dev set: supported 99%, unsupported 60%, contradicted 42%, partial 20%. It is agreement with lenient human labels, not the probability that a sentence is false. The judge almost always says it is highly certain, so stated certainty adds little; a real per-sentence probability needs the verdict token's logprob, which the gateway does not return today.
  • The prompt was revised twice on the dev split; the test split was run once. The test run used the direct route to the same vLLM server because the shared gateway was saturated at the time.
  • Small source sets (up to 24,000 characters) go to the judge whole, sources first, so the model server reuses its prefix cache across the sentences of one check; larger sets are narrowed per sentence with BM25.
Assemble it

Run this exact stack on your machine

Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.

grounding/assemble-prompt.md138 lines
# Assemble the Decosa grounding check on this machine

You are setting up a grounding checker: it takes a text (an answer, an email, a FAQ, a summary) and the sources it is
supposed to rest on, and judges every sentence as supported, partial, unsupported or contradicted, with the source span it
rests on. A claims gate turns that into pass, flag or block, and each check ends with a report signed by this box's own
key. Work step by step, show me each command before you run anything with `sudo`, and stop to ask if a check fails.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/grounding.zip (2 KB, 10 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py grounding` (the api image carries the same bundle under /app/rehearsal/grounding/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py grounding --bundle grounding.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the claims gate blocks the answer", "every claim sentence got a verdict (none errored)", "the battery-life sentence is contradicted or not backed"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## 0. Ground rules and licences
- Models: Qwen3.8-27B (Apache-2.0) as the judge. Optional CPU judge: `cross-encoder/nli-deberta-v3-base` (Apache-2.0).
  The checker is decosa-api (AGPL-3.0-or-later) and needs no GPU of its own.
- Drafts and sources stay on this machine. Bind every port to 127.0.0.1. The service keeps no text: nothing is written
  to disk and logs carry counts only. Keep it that way; do not add request logging.
- Be honest about what it does: "supported by the sources" is not "true". It reads only the sources it is given, and the
  judge is a language model that can be wrong in both directions.

## 1. Check the machine
1. `nvidia-smi`: one GPU with at least 32 GB for the judge (Qwen3.8-27B NVFP4 needs about 20 GB of weights plus KV
   cache; an RTX PRO 6000 96 GB or an RTX 5090 32 GB both work). Driver 570 or newer. Blackwell cards run NVFP4; on older
   cards use the FP8 weights.
2. No GPU? Use only the CPU judge (step 3, option B). It is weaker (see the numbers on the Stack tab) and makes no model
   call, so it has no model receipts.
3. `docker --version` and `docker compose version`. If Docker or the NVIDIA container toolkit is missing, install them
   from the official Docker and NVIDIA repositories after asking me, then run
   `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
4. Disk: about 30 GB free (Qwen3.8-27B NVFP4 about 20 GB, images, the 0.7 GB NLI model if used).

## 2. Images and weights
- `${DECOSA_REGISTRY}/decosa-api:<tag>` (**publishing soon**). If the pull fails, build from source:
  `git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host>` (access required), check out a release that contains
  `decosa_api/verticals/grounding/`, and build `docker/api/Dockerfile` as `decosa-api:local`. For the CPU judge, add
  a layer on top (the image runs as a non-root user, so install as root), build it with
  `docker build -t decosa-api:nli ./nli`, and use `decosa-api:nli` as the api image in option B (about 2.2 GB):
  ```dockerfile
  # ./nli/Dockerfile
  FROM decosa-api:local
  USER root
  RUN pip install --no-cache-dir torch --index-url https://download.pytorch.org/whl/cpu && pip install --no-cache-dir transformers sentencepiece
  USER decosa
  ```
- `vllm/vllm-openai:v0.29.0` for the judge; weights `nvidia/Qwen3.8-27B-NVFP4` (or `Qwen/Qwen3.8-27B-FP8` on a card
  without NVFP4).
- The CPU judge downloads `cross-encoder/nli-deberta-v3-base` from Hugging Face on first use; run once with network,
  then set `HF_HUB_OFFLINE=1`.

## 3. docker-compose.yml
Write this in `~/decosa/grounding/`.

Option A, the GPU judge (the tier the hosted demo runs):

```yaml
services:
  llm:
    image: vllm/vllm-openai:v0.29.0
    command: ["--model", "nvidia/Qwen3.8-27B-NVFP4", "--served-model-name", "qwen3.8-27b", "--max-model-len", "32768",
              "--enable-prefix-caching"]
    ports: ["127.0.0.1:8114:8000"]
    volumes: ["~/.cache/huggingface:/root/.cache/huggingface"]
    deploy: { resources: { reservations: { devices: [{ driver: nvidia, count: 1, capabilities: [gpu] }] } } }
    healthcheck: { test: ["CMD", "curl", "-fs", "http://localhost:8000/v1/models"], interval: 30s, retries: 20 }
  api:
    image: ${DECOSA_REGISTRY}/decosa-api:<tag>
    ports: ["127.0.0.1:8445:8445"]
    environment:
      DECOSA_HOST: 0.0.0.0
      DECOSA_PORT: "8445"
      DECOSA_DATA_DIR: /data
      DECOSA_LLM_ROUTE: direct
      DECOSA_LLM_URL: http://llm:8000/v1
      DECOSA_LLM_MODEL: qwen3.8-27b
      DECOSA_GROUNDING_MAX_CONCURRENT: "4"
      DECOSA_GROUNDING_WORKERS: "6"
    volumes: ["decosa-data:/data"]
    depends_on: { llm: { condition: service_healthy } }
    healthcheck: { test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://localhost:8445/grounding/info', timeout=4)"], interval: 30s, retries: 10 }
volumes:
  decosa-data:
```

The api keeps its state (keys, receipts, this box's signing key) in the named volume `decosa-data`, not in a
host folder: the image runs as an unprivileged user (uid 10001), and a host folder that Docker creates is owned by
root, which stops the api with `PermissionError: [Errno 13] Permission denied: '/data/keys.sqlite'`. Then start everything: `docker compose up -d`.

Prefix caching matters here: every sentence of one check sends the same sources first, so the server reuses that part.

Option B, CPU only: drop the `llm` service and `depends_on`, and add to `api.environment`
`DECOSA_GROUNDING_NLI: cross-encoder/nli-deberta-v3-base` and `HF_HOME: /data/hf`. Requests then send `"judge": "nli"`.

On the first start the api service creates this box's Ed25519 key in the `decosa-data` volume (`/data/attest/` in the api container, mode 0600). Back it up with
`docker compose cp api:/data/attest ./attest-backup` and keep that copy private.
Never print it. Every model call on the direct route gets a receipt signed with that key (status `attested`): an
attestation by me, the operator, not a proof of computation.

## 4. Smoke test
1. `curl -s localhost:8445/grounding/info | head -c 600` lists the verdicts, the limits and the judge.
2. Get a token: `T=$(curl -s -XPOST localhost:8445/demo/session -H 'content-type: application/json' -d '{"vertical":"grounding"}' | jq -r .token)`.
3. `curl -s localhost:8445/grounding/samples | jq '.[2]' > sample.json` (the thermostat support answer), then
   `jq '{text, sources, question}' sample.json | curl -s -XPOST localhost:8445/grounding/gate -H "authorization: Bearer $T" -H 'content-type: application/json' -d @- > gate.json`.
   Expect `decision: "block"`, the battery-life and 240 V sentences `contradicted`, and the reset sentence `supported`
   with evidence span `S1.3`. With option B the verdicts are coarser; report what you get.
4. Stream the same body to `/grounding/check` with `-H 'accept: text/event-stream' -N`: a `ready` event, then a `receipt`
   and a `sentence` event per sentence, then `report`, `budget` and `done`.
5. `jq '{report}' gate.json | curl -s -XPOST localhost:8445/grounding/verify -H 'content-type: application/json' -d @-`
   must show `valid_signature: true` and `signed_by_this_server: true`. Then change one verdict in the report and
   verify again: it must fail.
6. Time it: on our RTX PRO 6000 a check of four or five sentences took 0.9-3.0 s through a quiet gateway, and single
   judge calls averaged about 2 s with six in flight while the card was shared. Tell me what you measure.

## 5. Point the app at the local API
Set `NEXT_PUBLIC_DECOSA_API=http://127.0.0.1:8445` in the site's `.env.local`, or call the routes directly from the
code that sends emails or publishes pages: `POST /grounding/gate` before sending, and keep the signed report with the
sent item. Contract: `API_CONTRACT.md`, section "Grounding check and outbound claims gate".

Off by default. Joining serves other people's requests on this GPU; never do it on a box that holds confidential drafts.
If I ask for it, follow the provider guide at `/provide` on the site, and do not enable it without my explicit yes.
Technical detailsModels, where it runs, labels

In short

Last reviewed

What it is
Is each sentence supported by the sources? Per-sentence verdicts with the source span, a claims gate, and a signed report.
Who it's for
Teams that ship AI-written answers, emails or FAQs and want each sentence checked against its sources first.
Where it runs
Hosted or self-host
Key numbers
  • 0.593 / 0.556 Unsupported-sentence precision / recall, default gate (held out, n = 2069)
  • 0.966 Recall, strict gate (partial counts too) (held out, n = 2069)
  • 0.691 Response-level F1, default gate (held out, n = 300)
  • 17.2 s Median end-to-end run, hosted (QA sweep 2026-09-25)
All results, datasets and caveats
Models
Qwen3.8-27B
Where
Hosted or self-host
Checks
Receipt per verdict; signed report
Output
Structured data · Signed record or verdict
Data
Confidential business data
Hardware
1× 96 GB GPU
Licence
Permissive (Apache-2.0, MIT)

Questions people ask

What does the Decosa grounding check do?

The Decosa grounding check splits a draft into sentences and asks an open judge model about each one in its own call: supported, partial, unsupported or contradicted, with the source spans it rests on. A claims gate turns the verdicts into pass, flag or block before an email, FAQ or answer ships, and every check ends with an Ed25519-signed report of hashes and offsets. It reads only the sources you send, so supported is not the same as true.

How accurate is the grounding check at catching hallucinations?

On 2,069 held-out RAGTruth sentences (178 unsupported), the Decosa grounding check's default gate found unsupported sentences with precision 0.593 and recall 0.556 (F1 0.574). A strict gate that also counts partial reached recall 0.966 at precision 0.276, flagging 30.1% of sentences. Response-level F1 was 0.691. This is one English benchmark built from 2023-era model outputs, so measure the false-flag rate on your own domain before blocking without a human look.

Is a grounding check the same as a fact check?

No. The Decosa grounding check is a quality control, not a fact check: it compares each sentence only with the sources you provide, so a sentence copied from a wrong source passes. Verdicts come from a language model and can be wrong in both directions, as the published RAGTruth precision and recall show. Sources must be pasted text; a bare URL is refused because nothing is fetched.

Does the grounding check store my drafts or sources?

The hosted Decosa grounding check keeps no text: drafts and sources stay in memory for the request, and the signed report carries hashes and character offsets only. Inputs are up to 8,000 characters and 40 sentences, with up to 10 sources and 60,000 characters in total. For unpublished or confidential drafts, self-host; on your own box the judge runs on your vLLM and nothing leaves it.

Can I run the grounding check without a GPU?

Yes, with a large accuracy cost. The Lite tier of the Decosa grounding check uses an NLI cross-encoder on CPU; on the same RAGTruth test it scored F1 0.227 (precision 0.133, recall 0.792, flag rate 51.4%) and on the thermostat sample it missed one of two contradictions. The Standard tier runs Qwen3.8-27B on one GPU, measured on a 96 GB RTX PRO 6000, and is what the hosted API uses.

Does a pass prove a marketing claim is substantiated?

No. Where rules require substantiating marketing claims, such as the FTC's advertising substantiation policy in the US, a pass from the Decosa grounding check is a record that a check ran against the sources you chose, not proof that a claim is substantiated. It is not legal advice. A hosted check of a five-sentence answer against a one-page source cost about $0.002 on 2026-09-25.

Ask a question or leave feedbackWe read every message and publish useful answers
Questions & feedback

Ask about Grounding check

We read every message. Questions, comments and our answers show here once we have reviewed and approved them.

Loading questions…

This is a

Plain text. Please leave out personal, patient or client data.

Shown with your message if we publish it. Leave blank to post as “A visitor”.

Nothing appears here until we have read and approved it.