Check CSR numbers against tables
A QC report that points every number in the narrative at the table cell it reports and names the likely slip behind each mismatch.
Built on: Document reader, Evidence retrieval, Numeric grounding, Signed record
Loading the tool…
Use it your way
Use it from your codeThe hosted API with your key, and prompts to paste into a coding agent
Run it yourself, on request
- The same open models and app, on 1x RTX PRO 6000 (96 GB) for Qwen3.8-27B, plus about 10 GB for the embedder and reranker and about 6 GB for the document reader; the comparisons, report and record run on CPU.
- Data never leaves your machines, and there are no Decosa charges.
- One prompt for Claude Code or Codex assembles the whole stack.
- Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Get an API key
- Call the csr number-to-table verifier API from your own code in minutes.
- Every model answer carries a signed receipt.
- Synthetic, public or test data only: real confidential data belongs on your own hardware.
Build with it
Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.
- Base URL
- https://api.decosa.ai
- Auth
Authorization: Bearer $DECOSA_API_KEY(or a demo session token)- Tool id
- csr-number-verifier
Use the hosted API
# Decosa CSR number-to-table verifier: use the hosted API
You are wiring Decosa's CSR number check into this project (a medical-writing QC workflow, a document tool or a
script). It takes a clinical study report's narrative (sections 10 to 12, or any sections you name) with its in-text
tables and TLFs, as text and rows or as a PDF, and traces every number in the narrative to the table cell it reports.
It returns each number with its status and the cell it cites, flags mismatches, numbers with no source and derived
numbers that don't compute (with the likely slip), compares in-text tables with their source TLF, and signs a record.
Every model call has a signed receipt. Use only what is listed below; if you need something else, stop and ask me.
- Base URL: `https://api.decosa.ai`
- Health check: `GET https://api.decosa.ai/healthz`.
- **The hosted API takes invented or public reports only** (`"synthetic": true` is required). Unblinded study results
are confidential: use the self-hosted version for a real report. Never send a real sponsor's CSR here.
- This is a medical-writing QC aid, not a validated system and not a regulatory submission tool. The medical writer and
the QC reviewer decide what to change; the TLFs themselves are the biostatistician's to validate. An unflagged number
is not proven right.
## Auth: API key (or a demo session)
1. Preferred: an API key (`dk_…`) from "Get an API key" on the tool page. Keep it in `DECOSA_API_KEY`, never in code.
Send `Authorization: Bearer $DECOSA_API_KEY`.
2. Without a key: `POST https://api.decosa.ai/demo/session` with `{"vertical": "csr-number-verifier"}` returns `{"token", "expires_at", "budget"}`.
Demo sessions are limited per IP per hour and carry a generated-token budget. Over a limit you get HTTP 429 with
`Retry-After`; 402 when the budget left is too small; a demo token runs one check at a time (409 otherwise).
## Check a report
- `POST /csr/verify` (token). Body, one of:
- `{"csr": {"title"?: "...", "narrative": "<text with headings like 11.4.1 Primary Efficacy>" | [{"section", "heading", "text"}],
"tables": [{"id": "14.2.1", "title": "...", "rows": [["", "Placebo (N=206)", "Drug 40 mg (N=205)"], ["Responders, n (%)", "52 (25.2)", "71 (34.6)"]],
"header_rows"?: 1, "source"?: "14.2.1"}], "previous_tables"?: [...], "previous_cut"?: "interim cut 15 Jan 2026"}, "synthetic": true}`
(a table with `source` is an in-text table and is compared with that TLF; `previous_tables` are the same TLFs at an
earlier data cut, so a stale number can be named as such);
- `{"file_b64": "<PDF or page image, base64>", "media_type"?: "application/pdf", "synthetic": true}` (read by the document
reader block: born-digital text from the text layer, tables and scans from pixels);
- `{"sample": "zenavotide-301"}` (a bundled fictional report; `zenavotide-301-clean` has no errors; `-pdf` and `-scan`
are the same report as a PDF and as a scan).
- Options: `"sections": ["11", "12.2"]`, `"confirm": false` (skip the second reading), `"stream": true`.
- Limits: 400,000 characters of narrative, 120 tables of up to 400 rows and 20 columns, a 12 MB file of up to 40 pages.
- The JSON response has `results` (every number: `id`, `section`, `sentence`, `start`/`end` in the sentence, `text`,
`kind`, `status` one of traced, traced_by_value, derived_ok, mismatch, derived_wrong, untraceable, context; `cite`
with table, cell, row, column and value; `expected`; `slip` with type and detail; `derived`), `flags` (result ids),
`table_mismatches`, `counts`, `markdown` (the QC report), `index` (the retrieval index hash when a report has more than
four tables), `search_receipts`, `record`, `receipts` and, for a PDF, `read` (tables found and the document receipt).
- With `Accept: text/event-stream` (or `"stream": true`) the events are, for a PDF, `read_ready`, `read_page` per page and
`read_done`; then `ready`, `index`, a `receipt` per model call, a `paragraph` per paragraph with its numbers,
`tables_checked`, `summary`, `report`, `budget` and `done`.
- `POST /record/verify` (no token) `{"record": {...}}` re-checks the signed record.
- `GET /csr/info`, `GET /csr/samples`, `GET /csr/samples/{id}`, `GET /csr/samples/{id}/pdf` and `GET /attest/signing-key` need no token.
## Example (Python, `pip install httpx`)
```python
import base64, httpx, os, pathlib
API = "https://api.decosa.ai"
H = {"Authorization": f"Bearer {os.environ['DECOSA_API_KEY']}"}
pdf = base64.b64encode(pathlib.Path("invented_csr.pdf").read_bytes()).decode()
r = httpx.post(f"{API}/csr/verify", json={"file_b64": pdf, "synthetic": True, "sections": ["10", "11", "12"]},
headers=H, timeout=900)
r.raise_for_status()
js = r.json()
for x in js["results"]:
if x["status"] in ("mismatch", "derived_wrong", "untraceable"):
c = x.get("cite") or {}
print(f"{x['section']:>7} {x['text']:>8} {x['status']:<13} table {c.get('table')} {c.get('cell')} gives {x.get('expected')} {(x.get('slip') or {}).get('detail') or ''}")
pathlib.Path("number-qc.md").write_text(js["markdown"])
```
Run it yourself (containers)
On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.
# Decosa CSR number-to-table verifier: run it yourself (containers)
You are setting up Decosa's CSR number check on this machine, so an unblinded clinical study report never leaves it. It
reads the report (text and table rows, or a PDF or scan through the document reader), traces every number in the
narrative to the table cell it reports, compares and recomputes in code, flags what does not match, and writes a QC
report and a signed record. Nothing is sent to Decosa's hosted API. It is a medical-writing QC aid: the medical writer and
the QC reviewer decide.
Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.
Ask me before any command that needs sudo, and show me the command first.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/csr-number-verifier.zip (15 KB, 15 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py csr-number-verifier` (the api image carries the same bundle under /app/rehearsal/csr-number-verifier/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py csr-number-verifier --bundle csr-number-verifier.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the transposed mean age (75.5 for 57.5) is flagged", "the rounding slip in disposition (71.8% for 71.9%) is flagged", "the wrong N in the analysis sets (313 for 317) is flagged"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
instructions for this distribution (docs.docker.com/engine/install). Install the NVIDIA container toolkit and check
`docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
`mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
Read it. Keep the `llm` service (Qwen3.8-27B on vLLM) and the `api` service. For the `api` service set
`DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b`,
`DECOSA_CSR_SYNTHETIC_ONLY=0` and `DECOSA_DOCREADER_SYNTHETIC_ONLY=0` (so this box accepts real reports), and bind every
port to 127.0.0.1. Never set the gateway route on a box that holds a real report.
3. Retrieval (finds the right table in a long report): add the `retrieval` service from the decosa-api source
(`services/retrieval`, Qwen3-Embedding-0.6B and Qwen3-Reranker-4B, about 10 GB of GPU memory) and set
`DECOSA_RETRIEVAL_URL=http://retrieval:8499` on the api, or set `DECOSA_CSR_RETRIEVAL=bm25` for keyword search only.
4. PDFs and scans (optional): add the document reader (`services/docreader` plus PaddleOCR-VL-1.6 on vLLM, about 6 GB)
and set `DECOSA_DOCREADER_URL`. Without it, send the narrative as text and the tables as rows.
5. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check.
6. Check: `curl -fsS http://127.0.0.1:<PORT>/csr/info` shows `synthetic_only: false`; `GET /attest/signing-key` shows
this box's public key.
7. Smoke test: get a token with `POST /demo/session {"vertical":"csr-number-verifier"}` and send `{"sample": "zenavotide-301"}`
to `POST /csr/verify`. Expect flags in sections 10.1, 11.1, 11.2, 11.4.3, 12.2.1 and 12.3.1 (the six seeded errors)
and one in-text table mismatch (Table 11-1 against 14.2.1). Then `{"sample": "zenavotide-301-clean"}` should give no
flags. Send the first record to `POST /record/verify`: `ok` must be true.
8. Report back: the public key, the flag counts, and how long each run took.
Off by default. Joining as a provider serves other people's requests on this GPU; never do it on a box that holds a
sponsor's study data. If I ask for it later, follow the Provide page instead of improvising.
Run it on your own hardwareWhat it needs, and the prompt that sets it up
Run it on your own GPU
Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.
Hardware check
Check your own hardware- CPU only, 64 GB RAMDoesn't fit
Qwen3.8-27B (NVFP4) needs a GPU.
- GeForce RTX 4090lite tierRuns with a smaller tier
The standard tier does not fit: Needs about 35 GB of GPU memory at the smallest settings; 24 GB available. The lite tier fits with changes.
- GeForce RTX 5090lite tierRuns with a smaller tier
The standard tier does not fit: Needs about 43 GB of GPU memory at the smallest settings; 32 GB available. The lite tier fits with changes.
- 2x GeForce RTX 5090standard tierRuns
The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
- L40Slite tierRuns with a smaller tier
The standard tier does not fit: Needs about 48.6 GB of GPU memory at the smallest settings; 48 GB available. The lite tier fits with changes.
- H100 80 GB (SXM)standard tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.
- RTX PRO 6000 Blackwell 96 GBstandard tierRuns
The standard tier fits (72.6 of 96 GB).
- 2x RTX PRO 6000 Blackwell 96 GBstandard tierRuns
The standard tier fits (72.6 of 192 GB).
- Apple M3 Ultra (Mac Studio), 96 GBlite tierRuns with a smaller tier
The standard tier can't be checked: Evidence retrieval block: Qwen3-Embedding-0.6B + Qwen3-Reranker-4B (decosa-retrieval service) has no mapped Apple Silicon build The lite tier fits with changes.
- Apple M5 Max, 64 GBlite tierRuns with a smaller tier
The standard tier can't be checked: Evidence retrieval block: Qwen3-Embedding-0.6B + Qwen3-Reranker-4B (decosa-retrieval service) has no mapped Apple Silicon build The lite tier fits with changes.
Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.
On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.
- 1
Check the GPU, Docker and the NVIDIA Container Toolkit
The driver must see the GPU, and Docker must be able to pass it into a container.
nvidia-smi docker compose version docker run --rm --gpus all ubuntu nvidia-smi
- 2
Fetch the compose file
One file describes the API and the language model as services.
mkdir -p ~/decosa && cd ~/decosa curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml - 3
Pull and start
The first start downloads pinned model weights, tens of gigabytes.
docker compose pull docker compose up -d
- 4
Check health
Wait until the API reports ok with the language model loaded. Then point your app at the local base URL.
curl -fsS http://localhost:<PORT>/healthz # {"ok": true, "llm": true, ...} curl -fsS -X POST http://localhost:<PORT>/demo/session \ -H 'Content-Type: application/json' -d '{"vertical":"csr-number-verifier"}'
Set up with a coding agent, rehearse on mock data, then go private
- Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
- Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's
expected.json. Every check must print PASS. - Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
docker compose exec api python scripts/rehearse.py csr-number-verifier
Download the mock-data bundle (15 KB, 15 checks)expected.json
ZEN-395, an invented phase 3 psoriasis trial of an invented drug (zenavotide) with placebo and two doses: sections 10-12 of the narrative and eleven tables (nine TLFs, two in-text tables), plus the same TLFs at an earlier (interim) data cut. Seeded: transposed digits (75.5 for 57.5), a last-digit rounding slip (71.8% for 71.9%), a wrong N (313 for 317), a percentage on the wrong denominator (2.2% for 3.7%), the other arm’s count (72 for 14) and a percentage left over from the interim cut (53.5% for 58.3%), and one stale cell in in-text Table 11-1. The run must flag each seeded number (the other-arm count through its percentage, which then does not compute), keep the clean report clean, sign a record that verifies and fails when changed, and read the same report from a PDF.
What the rehearsal checks
- the transposed mean age (75.5 for 57.5) is flagged
- the rounding slip in disposition (71.8% for 71.9%) is flagged
- the wrong N in the analysis sets (313 for 317) is flagged
- the percentage left over from the interim cut (53.5% for 58.3%) is flagged
- the percentage on the wrong denominator (2.2% for 3.7%) is flagged
- the other-arm count in 11.4.3 is caught: its percentage no longer computes
- the stale cell in in-text Table 11-1 is found against its source TLF 14.2.1
- at least 70 numbers are checked
- the report says who decides
- the signed record verifies
- the record fails once its flag count is changed
- the clean report comes back with at most one flag
- the PDF is read into at least ten of its eleven tables
- reading the PDF, at least four of the seeded numbers are flagged
- every model call has a signed receipt (the document reader’s own receipt, dr-..., is a different format and is checked by the reader)
Licence: Synthetic: the drug, sponsor, trial and every number are invented (decosa_api/verticals/csr/synth.py, seed 72; scripts/csr_samples.py). Part of decosa-api, AGPL-3.0-or-later.
Prompt for your coding agent
# Decosa CSR number-to-table verifier: run it yourself (containers)
You are setting up Decosa's CSR number check on this machine, so an unblinded clinical study report never leaves it. It
reads the report (text and table rows, or a PDF or scan through the document reader), traces every number in the
narrative to the table cell it reports, compares and recomputes in code, flags what does not match, and writes a QC
report and a signed record. Nothing is sent to Decosa's hosted API. It is a medical-writing QC aid: the medical writer and
the QC reviewer decide.
Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.
Ask me before any command that needs sudo, and show me the command first.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/csr-number-verifier.zip (15 KB, 15 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py csr-number-verifier` (the api image carries the same bundle under /app/rehearsal/csr-number-verifier/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py csr-number-verifier --bundle csr-number-verifier.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the transposed mean age (75.5 for 57.5) is flagged", "the rounding slip in disposition (71.8% for 71.9%) is flagged", "the wrong N in the analysis sets (313 for 317) is flagged"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
instructions for this distribution (docs.docker.com/engine/install). Install the NVIDIA container toolkit and check
`docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
`mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
Read it. Keep the `llm` service (Qwen3.8-27B on vLLM) and the `api` service. For the `api` service set
`DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b`,
`DECOSA_CSR_SYNTHETIC_ONLY=0` and `DECOSA_DOCREADER_SYNTHETIC_ONLY=0` (so this box accepts real reports), and bind every
port to 127.0.0.1. Never set the gateway route on a box that holds a real report.
3. Retrieval (finds the right table in a long report): add the `retrieval` service from the decosa-api source
(`services/retrieval`, Qwen3-Embedding-0.6B and Qwen3-Reranker-4B, about 10 GB of GPU memory) and set
`DECOSA_RETRIEVAL_URL=http://retrieval:8499` on the api, or set `DECOSA_CSR_RETRIEVAL=bm25` for keyword search only.
4. PDFs and scans (optional): add the document reader (`services/docreader` plus PaddleOCR-VL-1.6 on vLLM, about 6 GB)
and set `DECOSA_DOCREADER_URL`. Without it, send the narrative as text and the tables as rows.
5. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check.
6. Check: `curl -fsS http://127.0.0.1:<PORT>/csr/info` shows `synthetic_only: false`; `GET /attest/signing-key` shows
this box's public key.
7. Smoke test: get a token with `POST /demo/session {"vertical":"csr-number-verifier"}` and send `{"sample": "zenavotide-301"}`
to `POST /csr/verify`. Expect flags in sections 10.1, 11.1, 11.2, 11.4.3, 12.2.1 and 12.3.1 (the six seeded errors)
and one in-text table mismatch (Table 11-1 against 14.2.1). Then `{"sample": "zenavotide-301-clean"}` should give no
flags. Send the first record to `POST /record/verify`: `ok` must be true.
8. Report back: the public key, the flag counts, and how long each run took.
Off by default. Joining as a provider serves other people's requests on this GPU; never do it on a box that holds a
sponsor's study data. If I ask for it later, follow the Provide page instead of improvising.
Help me customise for my hardware
Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.
GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page
Runs with a smaller tierCSR number-to-table verifier on GeForce RTX 5090: use the Lite · text and rows, keyword search tier
The standard tier does not fit: Needs about 43 GB of GPU memory at the smallest settings; 32 GB available. The lite tier fits with changes.
What this tool's stack says about this hardware:
- 1x RTX 5090 32 GB (does not fit): Estimate: Qwen3.8-27B NVFP4 (about 20 GB with a small KV cache) plus the reranker and reader is too tight; run the lite tier (text and rows, keyword search) or put the retrieval and reader services on a second card.
Lite · text and rows, keyword search: what changesuses estimates
- Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
Memory per component
- Finds the numbers, keeps the tables as typed...: decosa-api CSR verifier (decosa_api/verticals/csr), importing the document reader, the evidence retrieval block and the numeric-grounding block's rounding helpers. CPU. Runs on CPU (vram_gb 0 in stack.json).
- One call per paragraph: Qwen3.8-27B (NVFP4). ~57.6 GB (at least ~28 GB), weights 21.4 GB (from stack.json). Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)
Expected speed
Not measured.
Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.
Setup prompt for this hardware
The self-host prompt for CSR number-to-table verifier, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.
# Set up CSR number-to-table verifier on my hardware Fetch https://decosa.ai/prompts/csr-number-verifier-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied. ## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=csr-number-verifier) Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4). Quality tier: Lite · text and rows, keyword search (lite). Fit check: runs with changes, about 28 GB of 32 GB used; some memory numbers are estimates, not measurements. First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything. Use these components (the setup below describes the standard tier; change it to match): - Finds the numbers, keeps the tables as typed...: decosa-api CSR verifier (decosa_api/verticals/csr), importing the document reader, the evidence retrieval block and the numeric-grounding block's rounding helpers, CPU - One call per paragraph: Qwen3.8-27B (NVFP4) (nvidia/Qwen3.8-27B-NVFP4), 57.6 GB. Change: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions. GPU placement (set each service's device and its vLLM --gpu-memory-utilization to about the share shown): - GPU 0: Qwen3.8-27B (NVFP4) ~28 GB (88%); about 4 GB left During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed. The stack's own component list and compose layout: https://decosa.ai/prompts/csr-number-verifier-assemble.md
The proof
How we tested itEval results and end-to-end checks, hosted and self-hosted, with dates
Verified end to end
Hosted: verified 27 Sep 2026 · measured 27 Sep 2026: · p50 6.5 s · ~$0.004 per run · 6 receipts
Loading the nightly status…
Self-host: verified 27 Sep 2026 · fresh clone into a clean directory, api image from docker/api/Dockerfile, compose api with a named volume, direct route to the local Qwen3.8-27B, the running document reader (:8497) and retrieval (:8499) services, local signing; torn down after
Measured cost to run: about $0.087 per 100 pages (hosted, 27 Sep 2026, partly estimated). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.
The rehearsal bundle passed 15/15 in 62 s; every receipt attested. The reader and retrieval services were the running ones, not built from the compose file here.
Known limits (5)
- Hosted verification ran on the pre-release server (decosa-api the pre-release branch on our server, gateway route). The production API gets this tool when the branch merges.
- Measured on synthetic CSRs written by the same author as the prompts and on ClinicalTrials.gov results with a template narrative; not on a real sponsor CSR or against a QC reviewer's findings.
- Figures, listings and patient narratives are not read; RTF and SAS outputs must be exported to PDF or rows first.
- A PDF run reads up to 40 pages; split a longer report by section.
- An unflagged number is not proven right: 'matched by value only' means the value was found in one cell, not that the sentence was matched to it. A run where model calls failed is marked incomplete.
How it's builtThe steps, the models and what each one checks
Run it yourself, on request
- The same open models and app, on 1x RTX PRO 6000 (96 GB) for Qwen3.8-27B, plus about 10 GB for the embedder and reranker and about 6 GB for the document reader; the comparisons, report and record run on CPU.
- Data never leaves your machines, and there are no Decosa charges.
- One prompt for Claude Code or Codex assembles the whole stack.
- Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Get an API key
- Call the csr number-to-table verifier API from your own code in minutes.
- Every model answer carries a signed receipt.
- Synthetic, public or test data only: real confidential data belongs on your own hardware.
Every number in a clinical study report traced to the table cell it reports, checked in code, with the slips flagged.
For medical writing and regulatory operations teams at sponsors and CROs, and QC contractors who do the independent number check. Give it a clinical study report: the narrative of sections 10 to 12 with its in-text tables and TLFs, as text and table rows or as a PDF (born-digital or scanned, read by the document reader block). Code finds every number in the narrative and sets aside context (table references, doses, visit weeks, the 95% of a CI). For each paragraph, the evidence retrieval block picks the tables it most likely reports, after any it names, and Qwen3.8-27B points each number at the cell the sentence claims to report: the right arm, row and timepoint. Code then compares the value at the printed decimals, recomputes percentages, differences, totals and relative reductions, and names the likely slip for a mismatch: the other arm's value, a value from an earlier data cut, transposed digits, rounding, a percentage on the wrong denominator. In-text tables are compared with their source TLF cell by cell. You get a QC report and a signed record of every number's cell and verdict. A medical-writing QC aid: the medical writer and the QC reviewer decide.
- Deployment
- Self-host first
- Regulatory
- No regulation requires this check; it supports the manual number QC that sponsors and CROs already run on clinical study reports. Written 27 Sep 2026. ICH E3 (Structure and Content of Clinical Study Reports) defines the CSR's sections, including the efficacy (11) and safety (12) evaluations and the tables that support them (not re-read for this note). The FDA's draft guidance 'Considerations for the Use of Artificial Intelligence to Support Regulatory Decision-Making for Drug and Biological Products' (January 2025; https://www.fda.gov/regulatory-information/search-fda-guidance-documents/considerations-use-artificial-intelligence-support-regulatory-decision-making-drug-and-biological) is still a draft and is about AI used to produce information supporting regulatory decisions; a QC aid that only flags numbers for a person to check is not that, but a sponsor's own quality system decides how to qualify any tool it uses. The EMA reflection paper on the use of AI in the medicinal product lifecycle (adopted 30 Sep 2024; https://www.ema.europa.eu/en/use-artificial-intelligence-ai-medicinal-product-lifecycle-scientific-guideline) asks sponsors to take a risk-based approach and keep a human accountable. This tool is not a validated (GxP) system and not a regulatory submission tool: the medical writer and the QC reviewer decide what to change, and the TLFs remain the biostatistician's responsibility. Dates and scope from the Decosa wiki's sourced notes (page 33, S35 and S36, checked 26 Sep 2026); read the documents before relying on this.
Text description
A clinical study report (narrative, in-text tables, TLFs; text and rows or a PDF) goes to decosa-api. The document reader turns PDF pages into paragraphs and table cells. Code finds the numbers; retrieval picks the tables for each paragraph. Qwen3.8-27B points each number at the cell it claims to report. Code compares and recomputes, names the likely slip, and checks in-text tables against their TLF. The output is a QC report and a signed record; the medical writer and QC reviewer decide.
At a glance
- Data retention
- Nothing stored: the report lives in memory for the request. The signed record holds each number's status, cell and expected value, table hashes and receipt ids, never the report text; logs carry counts and timings only.
- What leaves the box
- Self-hosted on the direct route: nothing. The model, the retrieval service, the document reader and the checks run on the same machine. Hosted: model calls go through the Decosa API, and only invented or public reports are accepted.
- What it checks
- Every number in the narrative (counts, Ns, percentages, means, differences, CIs, hazard ratios, p-values) against the cell it reports, derived numbers recomputed, in-text tables against their source TLF, and the likely slip for each mismatch.
- What it is not
- A validated (GxP) system or a regulatory submission tool. It does not check the TLFs themselves, figures, listings or interpretation. The medical writer and the QC reviewer decide.
Pick the tier for the quality you need
Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.
Lite
text and rows, keyword search
Send the narrative as text and the tables as rows (exported from the TLF outputs), and run with DECOSA_CSR_RETRIEVAL=bm25: no retrieval service and no document reader, so only the language model needs a GPU. Same pointer, checks and record.
- Models
- decosa-api CSR verifier (decosa_api/verticals/csr), importing the document reader, the evidence retrieval block and the numeric-grounding block's rounding helpers
- Qwen3.8-27B (NVFP4)
- Hardware
- 1x GPU for Qwen3.8-27B
- Quality evidence
- Planted number errors caught, held-out synthetic CSRs, keyword search93 / 96 (96.9%), 1 false flag in 899 clean numbersdecosa-api docs/evals/csr-number-verifier.md, test split, BM25 ablation
- Latency
- measured: about a minute per hundred pages on the held-out synthetic set (BM25 only, shared gateway)
- Verification
- Proof: strongSelf-host onlyEvery model call is receipted; searches are BM25 in code, so there are no embed or rerank attestations.
- In the hosted demo
Standard
retrieval + document reader + the model (hosted demo)
The retrieval block picks the tables per paragraph, the document reader reads PDFs and scans, Qwen3.8-27B points each number at its cell and code checks everything. This is what the hosted demo runs, on invented or public reports only.
- Models
- decosa-api CSR verifier (decosa_api/verticals/csr), importing the document reader, the evidence retrieval block and the numeric-grounding block's rounding helpers
- Qwen3.8-27B (NVFP4)
- Evidence retrieval block: Qwen3-Embedding-0.6B + Qwen3-Reranker-4B (decosa-retrieval service)
- Document reader block: Docling 2.130 (Heron layout) + PaddleOCR-VL-1.6 (0.9B)
- Hardware
- 1x RTX PRO 6000 96 GB (measured on shared cards)
- Quality evidence
- Planted number errors caught, held-out synthetic CSRs (template narrative)92 / 96 (95.8%)decosa-api docs/evals/csr-number-verifier.md, test split
- Planted number errors caught, held-out synthetic CSRs rewritten by a second writer (Qwen paraphrase, numbers kept)93 / 96 (96.9%)decosa-api docs/evals/csr-number-verifier.md, test-para split
- Planted number errors caught, 30 ClinicalTrials.gov trials laid out as CSR tables170 / 178 (95.5%)decosa-api docs/evals/csr-number-verifier.md, ctgov split
- False flags on clean reports (per 100 numbers checked)0 per 100 (0 / 899) on held-out synthetic; 0.11 (1 / 899) rewritten; 2.04 (19 / 931) on ClinicalTrials.gov tablesdecosa-api docs/evals/csr-number-verifier.md
- Traced numbers citing the right cell99.1% (884 numbers, held-out synthetic); 98.2% (901, ClinicalTrials.gov)decosa-api docs/evals/csr-number-verifier.md
- Time per 100 pages129 s from JSON; 300 s from a born-digital PDF, 493 s from a scan (reading included), shared gatewaydecosa-api docs/evals/csr-number-verifier.md
- Latency
- measured: a couple of minutes per hundred pages from JSON on the shared gateway (longer under heavier load); under a minute for the sample and its scan; seconds for the smoke check
- Verification
- Proof: strongEvery Qwen3.8 call is a separate gateway call with a gateway-signed receipt; every search has a signed search receipt; the PDF read has a signed document receipt; the record lists them all.
Also runs on
- Number-consistency checker tier (CPU)Number-consistency checker (decosa_api.checkers.numbers, on a pre-release build)not builtFor numbers the pointer leaves untraced: a small encoder trained on planted number errors reads the sentence against the table; it can mark a number supported but never overrules a code mismatch. The hook is built; the checker is on an unmerged branch and not measured with this tool. Hardware: CPU, in process.
We host these ourselves when needed: small models get more of our own compute unless we detect a shortage, so they need no community providers.
Every model in the stack
| Model | Tiers | Params · VRAM | Verification | Details |
|---|---|---|---|---|
Finds the numbers, keeps the tables as typed cells, compares and recomputes in code, names the likely slip, checks in-text tables against their TLF, writes the QC report and the signed record (no model; CPU)decosa-api CSR verifier (decosa_api/verticals/csr), importing the document reader, the evidence retrieval block and the numeric-grounding block's rounding helpers 0 GBProof: partial | LiteStandard | 0 GB | Proof: partial | |
| ||||
One call per paragraph: which cell each number claims to report (by arm, row and timepoint), or which cells a derived number is computed from; a second look for numbers left unplaced; a second reading (values hidden) between two neighbouring cellsQwen3.8-27B (NVFP4)nvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab) 27.8B · 20 GBProof: strongIn the hosted demo | LiteStandard | 27.8B · 20 GB | Proof: strongIn the hosted demo | |
| ||||
Finds the tables a paragraph most likely reports in a long report (after the tables it names), with a signed receipt per searchEvidence retrieval block: Qwen3-Embedding-0.6B + Qwen3-Reranker-4B (decosa-retrieval service)Qwen/Qwen3-Reranker-4B on Hugging Face (opens in a new tab) 0.6B + 4B · 9 GBProof: partialIn the hosted demo | Standard | 0.6B + 4B · 9 GB | Proof: partialIn the hosted demo | |
| ||||
PDFs and scans: finds and orders the regions of each page (Docling Heron layout) and reads tables as cells with spans (PaddleOCR-VL-1.6); born-digital text comes from the PDF's text layerDocument reader block: Docling 2.130 (Heron layout) + PaddleOCR-VL-1.6 (0.9B)PaddlePaddle/PaddleOCR-VL-1.6 on Hugging Face (opens in a new tab) 0.9B · 6 GBProof: partialIn the hosted demo | Standard | 0.9B · 6 GB | Proof: partialIn the hosted demo | |
| ||||
Optional tier for numbers the pointer leaves untraced: a small encoder trained on planted number errors reads the sentence against the tableNumber-consistency checker (decosa_api.checkers.numbers, on a pre-release build) 0 GBProof: partial | Alternate | 0 GB | Proof: partial | |
| ||||
Tools, services and hardware
Tools
- ClinicalTrials.gov results (API v2) (opens in a new tab)ClinicalTrials.gov terms (updated 31 Jan 2023): attribute ClinicalTrials.gov, show the processing date, state modifications
Eval only: 30 completed phase 3 trials' posted results laid out as CSR tables (percentages computed from the posted counts), with a template narrative; data retrieved 27 Sep 2026.
- Synthetic CSRs of fictional trialsAGPL-3.0-or-later (decosa-api, decosa_api/verticals/csr/synth.py)
The samples and the dev and test sets: invented drugs, sponsors and numbers, with an earlier data cut and six kinds of planted number error.
- EMA clinical data publication (Policy 0070) (opens in a new tab)Terms of use allow general information and non-commercial research only (EMA/144064/2019, Annex 1 and 2)
Not used: published CSRs there may not be used commercially, so they are not in the demo or the eval.
Services
- decosa-api:8445
${DECOSA_REGISTRY}/decosa-api:<tag>GET /csr/info, /csr/samples; POST /csr/verify (SSE or JSON); POST /record/verify. Keeps no report text.
- decosa-retrieval:8499
built from services/retrieval (no published image yet)Embedder and reranker for the evidence retrieval block, on the same GPU box.
- decosa-docreader:8497
built from services/docreader (no published image yet), with PaddleOCR-VL-1.6 on vLLMReads PDFs and scans into paragraphs and table cells with page and box.
- vLLM (model):8114
vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1Qwen3.8-27B NVFP4 behind our gateway (hosted) or called directly (self-host).
Hardware
- 1x RTX PRO 6000 Blackwell 96 GB Fits
Measured on our server: the eval and the hosted demo ran Qwen3.8-27B through the shared gateway, with the retrieval service (about 10 GB) and the document reader (about 6 GB) on another GPU.
- 1x RTX 5090 32 GB Does not fit
Estimate: Qwen3.8-27B NVFP4 (about 20 GB with a small KV cache) plus the reranker and reader is too tight; run the lite tier (text and rows, keyword search) or put the retrieval and reader services on a second card.
- CPU only Does not fit
The model needs a GPU. The comparisons, report and record run on CPU.
Latency per lane
- the seeded 12-page sample as text and rows (78 numbers, 11 tables), hosted gateway route25.0 s
Measuredmeasured on our server 2026-09-27 (pre-release server, gateway shared with other workloads): 25.0 s seeded, 27.1 s clean; console run from the browser 36 s
- the same report as a 12-page scanned PDF (read from pixels, then checked)57.6 s
Measuredmeasured on our server 2026-09-27 (pre-release server): 57.6 s including the document reader
Notes
- Every comparison and recomputation is code: the model only says which cell a sentence speaks of. A wrong number it points at a cell that does not hold it is flagged; one it points at a cell that happens to hold it (usually the other arm's) passes, which is where most misses come from.
- Held out: 93 of 96 planted errors caught with no false flags in 899 clean numbers (synthetic), 93 of 96 after a second writer rewrote the prose, 170 of 178 on ClinicalTrials.gov tables (2 false flags per 100 numbers); the right cell cited 98-99% of the time; 129 s per 100 pages from JSON, 300-493 s from PDF or scan.
- A QC aid: it does not validate the TLFs, read figures or listings, or judge interpretation; an unflagged number is not proven right, and the medical writer and QC reviewer decide.
Run this exact stack on your machine
Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.
# Assemble the Decosa CSR number-to-table verifier on this machine
You are setting up a self-hosted number check for clinical study reports on this Linux machine, for a medical writing or
regulatory operations team at a sponsor or CRO. It reads a CSR (the narrative of sections 10 to 12 with its in-text
tables and TLFs; as text and table rows, or as a PDF or scan) and returns:
- every number in the narrative with the table cell it reports (table, row, column, value), checked in code at the
printed decimals, with percentages, differences, totals and relative reductions recomputed;
- flags for numbers that don't match their cell, numbers with no source and derived numbers that don't compute, each
with what the table gives and the likely slip (the other arm's value, a value from an earlier data cut, transposed
digits, rounding, a wrong denominator), and in-text tables compared with their source TLF;
- a Markdown QC report and a signed, hash-chained record of every number's cell and verdict and every receipt id.
Work step by step, show me each command before running anything that needs sudo, and stop if a check fails.
**Before anything else, remind me:**
- Unblinded study results are confidential. On this box nothing leaves the machine: the language model, the retrieval
models, the document reader and the checks all run here. Never point it at a hosted gateway while it holds a real report.
- This is a medical-writing QC aid, not a validated (GxP) system and not a regulatory submission tool. The medical writer
and the QC reviewer decide what to change; the TLFs are the biostatistician's to validate. An unflagged number is not
proven right.
Repeat both points in your final summary.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/csr-number-verifier.zip (15 KB, 15 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py csr-number-verifier` (the api image carries the same bundle under /app/rehearsal/csr-number-verifier/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py csr-number-verifier --bundle csr-number-verifier.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the transposed mean age (75.5 for 57.5) is flagged", "the rounding slip in disposition (71.8% for 71.9%) is flagged", "the wrong N in the analysis sets (313 for 317) is flagged"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## What you are building
| service | image | model | port |
|---|---|---|---|
| `llm` | `${DECOSA_REGISTRY}/decosa-llm:0.1.0` (vLLM 0.29.0, `vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1`) | `nvidia/Qwen3.8-27B-NVFP4` @ `482ca0f3832238542f8f5295dde86b5f22711d80`, Apache-2.0 | internal 8000 |
| `retrieval` | built from `decosa-api/services/retrieval` | `Qwen/Qwen3-Embedding-0.6B` @ `97b0c614be4d77ee51c0cef4e5f07c00f9eb65b3` + `Qwen/Qwen3-Reranker-4B` @ `22e683669bc0f0bd69640a1354a6d0aebcfeede5`, Apache-2.0 | internal 8499 |
| `parser` | `vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1` | `PaddlePaddle/PaddleOCR-VL-1.6` @ `c5630abae1d940eafe0697512a0325494b02ab42`, Apache-2.0 | internal 8000 |
| `docreader` | built from `decosa-api/services/docreader` | Docling 2.130 (MIT) with `docling-project/docling-layout-heron` @ `8f39ad3c0b4c58e9c2d2c84a38465abf757272d8` (Apache-2.0) | internal 8497 |
| `api` | `${DECOSA_REGISTRY}/decosa-api:0.1.0` (no GPU) | none | `127.0.0.1:8445` |
Only need text and table rows (no PDFs)? Leave out `parser` and `docreader` and their `depends_on` entry, and drop
`DECOSA_DOCREADER_URL`.
## 1. Check the GPU, driver and Docker
1. Run `nvidia-smi`. I need one NVIDIA GPU with at least 64 GB (the model about 20 GB plus KV cache, the retrieval models
about 10 GB, the document reader about 6 GB) and driver 580 or newer.
- Blackwell (RTX PRO 6000, B200): use the defaults below (NVFP4). This is the measured setup.
- Hopper (H100/H200): set `LLM_MODEL=Qwen/Qwen3.8-27B-FP8`, `LLM_REVISION=main` and `LLM_GPU_UTIL=0.65`. Not measured.
- Smaller cards: tell me. Use keyword search (step 3) and leave the document reader out or put it on a second card
(`DECOSA_GPU` per service).
- `PARSER_GPU_UTIL` is a share of the whole card: 0.04 is about 3.8 GiB of a 96 GB card; raise it on a smaller one.
2. Check `docker --version`, `docker compose version` and `docker run --rm --gpus all ubuntu nvidia-smi`. If Docker or the
NVIDIA Container Toolkit is missing, install them from the official Docker and NVIDIA repositories, run
`sudo nvidia-ctk runtime configure --runtime=docker` and restart Docker.
3. Confirm about 90 GB of free disk.
## 2. Get the images and the source
The images are **on request** while self-host is in early access: ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry (set it as `DECOSA_REGISTRY`), pull access and the compose file. Try `docker pull ${DECOSA_REGISTRY}/decosa-{llm,api}:0.1.0`. If a pull
fails, build from source once `decosa-api` is published: clone it into `~/decosa-csr/decosa-api`, run
`docker build -f docker/api/Dockerfile -t ${DECOSA_REGISTRY}/decosa-api:0.1.0 .`, and `docker compose build llm` from
its compose file. If neither works, stop and tell me. The `retrieval` and `docreader` services are always built from the
clone, so clone it either way.
## 3. Write the compose file
Create `~/decosa-csr/.env`:
```bash
DECOSA_TAG=0.1.0
DECOSA_GPU=0
LLM_MODEL=nvidia/Qwen3.8-27B-NVFP4
LLM_REVISION=482ca0f3832238542f8f5295dde86b5f22711d80
LLM_MAX_LEN=65536
LLM_GPU_UTIL=0.62
PARSER_GPU_UTIL=0.04
DECOSA_SIGNER_NAME="<who signs these records, e.g. Example Pharma medical writing QC>"
```
Create `~/decosa-csr/docker-compose.yml` with exactly these services:
```yaml
name: decosa-csr
x-health: &health
interval: 15s
timeout: 5s
retries: 5
services:
llm:
image: ${DECOSA_REGISTRY}/decosa-llm:${DECOSA_TAG}
deploy: { resources: { reservations: { devices: [ { driver: nvidia, device_ids: ["${DECOSA_GPU:-0}"], capabilities: [gpu] } ] } } }
ipc: host
restart: unless-stopped
volumes: [hf-cache:/root/.cache/huggingface]
command: ["${LLM_MODEL}", "--revision", "${LLM_REVISION}", "--served-model-name", "qwen3.8-27b",
"--language-model-only", "--max-model-len", "${LLM_MAX_LEN}", "--gpu-memory-utilization", "${LLM_GPU_UTIL}",
"--max-num-seqs", "16", "--kv-cache-dtype", "fp8_e4m3", "--speculative-config", '{"method":"mtp","num_speculative_tokens":3}',
"--seed", "0", "--enable-force-include-usage", "--disable-uvicorn-access-log", "--host", "0.0.0.0", "--port", "8000"]
healthcheck: { <<: *health, test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8000/health', timeout=4)"], start_period: 900s }
retrieval:
build:
context: ./decosa-api/services/retrieval
dockerfile_inline: |
FROM python:3.12-slim
WORKDIR /app
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt --extra-index-url https://download.pytorch.org/whl/cu130
COPY *.py .
CMD ["python", "server.py"]
deploy: { resources: { reservations: { devices: [ { driver: nvidia, device_ids: ["${DECOSA_GPU:-0}"], capabilities: [gpu] } ] } } }
restart: unless-stopped
environment:
RETRIEVAL_HOST: 0.0.0.0
RETRIEVAL_PORT: "8499"
RETRIEVAL_EMBED: qwen3-emb-0.6b
RETRIEVAL_RERANK: qwen3-rr-4b
HF_HOME: /root/.cache/huggingface
volumes: [hf-cache:/root/.cache/huggingface]
healthcheck: { <<: *health, test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8499/health', timeout=4)"], start_period: 600s }
parser:
image: vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1
deploy: { resources: { reservations: { devices: [ { driver: nvidia, device_ids: ["${DECOSA_GPU:-0}"], capabilities: [gpu] } ] } } }
ipc: host
restart: unless-stopped
volumes: [hf-cache:/root/.cache/huggingface]
command: ["PaddlePaddle/PaddleOCR-VL-1.6", "--revision", "c5630abae1d940eafe0697512a0325494b02ab42", "--served-model-name", "PaddlePaddle/PaddleOCR-VL-1.6",
"--trust-remote-code", "--max-model-len", "8192", "--gpu-memory-utilization", "${PARSER_GPU_UTIL:-0.04}", "--max-num-seqs", "16",
"--no-enable-prefix-caching", "--mm-processor-cache-gb", "0", "--generation-config", "vllm", "--host", "0.0.0.0", "--port", "8000"]
healthcheck: { <<: *health, test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8000/health', timeout=4)"], start_period: 600s }
docreader:
build: { context: ./decosa-api/services/docreader }
deploy: { resources: { reservations: { devices: [ { driver: nvidia, device_ids: ["${DECOSA_GPU:-0}"], capabilities: [gpu] } ] } } }
restart: unless-stopped
depends_on: { parser: { condition: service_healthy } }
environment:
DOCREADER_PARSER: paddle-openai
DOCREADER_PARSER_URL: http://parser:8000/v1
DOCREADER_PARSER_MODEL: PaddlePaddle/PaddleOCR-VL-1.6
volumes: [hf-cache:/models]
healthcheck: { <<: *health, test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8497/health', timeout=4)"], start_period: 300s }
api:
image: ${DECOSA_REGISTRY}/decosa-api:${DECOSA_TAG}
restart: unless-stopped
depends_on: { llm: { condition: service_healthy }, retrieval: { condition: service_healthy }, docreader: { condition: service_healthy } }
environment:
DECOSA_LLM_ROUTE: direct # local model only; receipts are signed by this box's key ("attested")
DECOSA_LLM_URL: http://llm:8000/v1
DECOSA_LLM_MODEL: qwen3.8-27b
DECOSA_LOCAL_SIGNING: "on" # Ed25519 key created at /data/attest/ed25519.pem on first start
DECOSA_SIGNER_NAME: ${DECOSA_SIGNER_NAME}
DECOSA_RETRIEVAL_URL: http://retrieval:8499
DECOSA_DOCREADER_URL: http://docreader:8497
DECOSA_CSR_SYNTHETIC_ONLY: "0" # this box accepts real reports
DECOSA_DOCREADER_SYNTHETIC_ONLY: "0"
DECOSA_CSR_MAX_PAGES: "40"
DECOSA_SESSIONS_PER_IP_HOUR: "1000"
DECOSA_BUDGET_LLM_TOKENS: "400000" # per session; a 12-page report uses a few thousand generated tokens
DECOSA_SESSION_TTL_S: "28800"
DECOSA_CORS_ORIGIN_REGEX: '^https?://(localhost|127\.0\.0\.1)(:\d+)?$$'
ports: ["127.0.0.1:8445:8445"]
volumes: [decosa-data:/data]
healthcheck: { <<: *health, test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8445/healthz', timeout=4)"], start_period: 20s }
volumes: { hf-cache: {}, decosa-data: {} }
```
The retrieval, parser and docreader services download their models into the shared `hf-cache` volume on first start.
**No room for the retrieval models?** Delete the `retrieval` service and its `depends_on` entry, drop
`DECOSA_RETRIEVAL_URL`, and add `DECOSA_CSR_RETRIEVAL: bm25` to the api. Tables are then found by keyword search only
(explicit "Table 14.2.1" references in the text are always used first); the comparisons and the record are unchanged.
Use the named volumes exactly as written: a root-owned host bind mount makes the API fail on `/data/keys.sqlite`.
Run `docker compose up -d --build`, then poll `docker compose ps` until every service is healthy (the LLM takes 5-10
minutes the first time). `curl -s localhost:8445/csr/info | jq '{synthetic_only, blocks}'` should show
`synthetic_only: false`, and `curl -s localhost:8445/docreader/info | jq .service` should show the reader reachable.
## 4. Smoke test on the bundled fictional report
```bash
API=localhost:8445
TOKEN=$(curl -s $API/demo/session -H 'content-type: application/json' -d '{"vertical":"csr-number-verifier"}' | jq -r .token)
curl -s $API/csr/verify -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' -d '{"sample":"zenavotide-301"}' > /tmp/csr.json
jq -r '.results[] | select(.status=="mismatch" or .status=="derived_wrong" or .status=="untraceable") | "\(.section)\t\(.text)\t\(.status)\t\(.expected)\t\(.slip.detail // "")"' /tmp/csr.json
jq '.table_mismatches | length' /tmp/csr.json
jq '{record}' /tmp/csr.json | curl -s $API/record/verify -H 'content-type: application/json' -d @- | jq '{ok, summary}'
curl -s $API/csr/verify -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' -d '{"sample":"zenavotide-301-clean"}' | jq '.counts'
curl -s $API/csr/verify -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' -d '{"sample":"zenavotide-301-scan"}' | jq '.counts, .read.pages'
```
Pass if the seeded report gives flags in sections 10.1, 11.1, 11.2, 11.4.3, 12.2.1 and 12.3.1 (the six seeded errors:
rounding, a wrong N, transposed digits, the other arm's value, a stale interim value and a wrong denominator) and at
least one in-text table mismatch (Table 11-1 against 14.2.1); every receipt is `attested`; the record verifies; the clean
report gives no flags; and the scan is read (12 pages) with the same kinds of flags.
## 5. Check a real report
As a PDF (born-digital or scanned, up to 40 pages per run; split a longer report by section):
```bash
python3 - <<'PY' > req.json
import base64, json, pathlib
print(json.dumps({"file_b64": base64.b64encode(pathlib.Path("csr.pdf").read_bytes()).decode(), "sections": ["10", "11", "12"]}))
PY
curl -s $API/csr/verify -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' -d @req.json > qc.json
jq -r .markdown qc.json > number-qc.md
jq .record qc.json > number-qc-record.json
```
Or as text and rows, which is more exact than reading a PDF: `{"csr": {"narrative": "<sections 10-12 as text, headings like
11.4.1 Primary Efficacy>", "tables": [{"id": "14.2.1", "title": "...", "rows": [[...], ...]}], "previous_tables": [...the
interim TLFs, optional...]}}`. Export TLF outputs (RTF, SAS) to PDF or rows first; RTF is not read directly.
## 6. Point tools at the local API
- Base URL `http://localhost:8445`; for a web app set `NEXT_PUBLIC_DECOSA_API=http://localhost:8445`.
- `POST /csr/verify` returns JSON, or Server-Sent Events with `Accept: text/event-stream`. `GET /csr/info` lists what it
checks and what it does not.
- Keep the QC report and the record with the document's QC file. Anyone can re-check the record with
`POST /record/verify` against the key at `GET /attest/signing-key`.
- Keep the API on 127.0.0.1; for other users, put a TLS reverse proxy with authentication in front.
## 7. Keep the direct route
`DECOSA_LLM_ROUTE=gateway` would send the report's sentences and tables to the hosted Decosa
API. Leave it off for real reports.
Finish with a summary: what is running, the health output, the smoke-test results, and the reminders above.Rules and regulations it checks againstDated, linked to the primary source; not legal advice
Regulation watch
Loading the watch status…
3 laws, rules and guidance pages cited; 2 watched nightly at the primary source. A change marks this page for a human re-check; nothing is edited automatically. What we cite and how it is watched
Technical detailsModels, where it runs, labels
In short
Last reviewed
- What it is
- Every number in a clinical study report traced to the table cell it reports, checked in code, with the slips flagged.
- Who it's for
- Medical writing and regulatory operations teams at sponsors and CROs, and medical-writing QC contractors.
- Where it runs
- Self-host for a real report; the hosted demo takes invented or public reports only
- Key numbers
- 170 / 178 Planted number errors caught, ClinicalTrials.gov tables (held out, n = 178)
- 92 / 96 Planted number errors caught (test split, n = 96)
- 93 / 96 Planted number errors caught, second writer (test split, n = 96)
- 6.5 s Median end-to-end run, hosted (QA sweep 2026-09-27)
- Models
- Qwen3.8-27B (points each number at its cell) · Qwen3-Embedding + Qwen3-Reranker (finding the right table in a long report) · PaddleOCR-VL + Docling layout (reading PDFs and scans)
- Where
- Self-host for a real report; the hosted demo takes invented or public reports only
- Checks
- Receipt per model call; signed search receipts with the index hash; every comparison and recomputation done in code; signed hash-chained record of each number's cell and verdict
- Industry
- Science and research · Healthcare
- Runs
- Self-host
- Output
- Signed record or verdict · Structured data
- Data
- Confidential business data
- Hardware
- 1× 96 GB GPU
- Licence
- Permissive (Apache-2.0, MIT)
- Runs in
- Self-host
- Built from
- Document reader · Evidence retrieval · Numeric grounding · Signed record
Questions people ask
Does the report leave our network?
Not when you self-host: the model, the retrieval service, the document reader and the checks all run on your own GPU box. The hosted demo accepts invented or public reports only and keeps nothing.
Can the model decide a wrong number is right?
It does no arithmetic and gives no verdict: it only says which cell a sentence speaks of, and code compares the value at the printed decimals and recomputes derived numbers. It can still let a slip through by pointing at a cell that happens to hold the written value, usually the other arm's: 3 of 24 such slips were missed in the held-out test.
Which errors does it look for?
Numbers that don't match their cell (transposed digits, the other arm's value, a wrong N, rounding slips, values left over from an earlier data cut), percentages on the wrong denominator, derived numbers that don't compute, numbers with no source, and in-text tables that differ from their TLF.
Does it read PDFs and scanned TLFs?
Yes, through the document reader block: born-digital text comes from the PDF's text layer, and tables and scans are read from the page pixels, up to 40 pages per run. Sending the tables as rows is more exact.
Is it a validated system for regulatory submissions?
No. It is a medical-writing QC aid: it finds and cites, and the medical writer and the QC reviewer decide what to change. It does not validate the TLFs.
Ask a question or leave feedbackWe read every message and publish useful answers
Ask about CSR number-to-table verifier
We read every message. Questions, comments and our answers show here once we have reviewed and approved them.
Loading questions…