Compare safety sections across labels
A list of every difference in the safety sections across the label documents and translations, quoted from each and marked likely error or deliberate.
Built on: Language pack, Document reader, Signed record
Loading the tool…
Use it your way
Use it from your codeThe hosted API with your key, and prompts to paste into a coding agent
Run it yourself, on request
- The same open models and app, on 1x RTX PRO 6000 (96 GB): Qwen3.8-27B NVFP4 plus Hy-MT2-7B (about 18 GB) and the document reader (about 6 GB); sections, number checks and the record run on CPU.
- Data never leaves your machines, and there are no Decosa charges.
- One prompt for Claude Code or Codex assembles the whole stack.
- Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Get an API key
- Call the label consistency across pi, smpc, ccds and carton API from your own code in minutes.
- Every model answer carries a signed receipt.
- Synthetic, public or test data only: real confidential data belongs on your own hardware.
Build with it
Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.
- Base URL
- https://api.decosa.ai
- Auth
Authorization: Bearer $DECOSA_API_KEY(or a demo session token)- Tool id
- label-consistency-check
Use the hosted API
# Decosa label consistency check: use the hosted API
You are wiring Decosa's label consistency check into this project (a labelling QC workflow, a document tool or a script).
It takes the label documents for one medicine (the company core data sheet (CCDS), the US Prescribing Information, the EU
SmPC and package leaflet in several languages, and carton text or a carton scan), aligns their sections (indications,
dosing, contraindications, warnings, adverse reactions, storage, strengths) and returns every difference with a quote
from each document and a triage: likely error, likely deliberate or unclear. Every model call has a signed receipt and
the run ends in a signed record. Use only what is listed below; if you need something else, stop and ask me.
- Base URL: `https://api.decosa.ai`
- Health check: `GET https://api.decosa.ai/healthz`.
- **The hosted API takes invented or public labels only** (`"synthetic": true` is required). Unapproved labelling is
confidential: use the self-hosted version for real labels. Never send a company's draft labelling here.
- It is a labelling QC aid, not a regulatory decision: the triage is a model judgment, and regulatory affairs decides
what is deliberate and what changes.
## Auth: API key (or a demo session)
1. Preferred: an API key (`dk_…`) from "Get an API key" on the tool page. Keep it in `DECOSA_API_KEY`, never in code.
Send `Authorization: Bearer $DECOSA_API_KEY`.
2. Without a key: `POST https://api.decosa.ai/demo/session` with `{"vertical": "label-consistency-check"}` returns `{"token", "expires_at", "budget"}`.
Demo sessions are limited per IP per hour and carry a generated-token budget: 429 with `Retry-After` over a limit, 402
when the budget left is too small, 409 when a demo token already has a run going.
## Check a label set
- `POST /label/check` (token). Body, one of:
- `{"documents": [{"role": "ccds", "text": "..."}, {"role": "us_pi", "text": "..."}, {"role": "eu_smpc", "lang": "en", "text": "..."},
{"role": "eu_smpc", "lang": "de", "text": "..."}, {"role": "eu_pl", "text": "..."}, {"role": "carton", "file_b64": "<PNG, JPEG or PDF>", "media_type": "image/png"}],
"declared_deviations": ["US PI: indication X not approved in the US.", "..."], "synthetic": true}`
(roles: ccds, us_pi, eu_smpc, eu_pl, carton; one per role and language; keep section headings in the text such as
"4.3 Contraindications" or "5 WARNINGS AND PRECAUTIONS");
- `{"sample_id": "norvexa-drift"}` (a bundled invented medicine; also `norvexa-clean`, `norvexa-quick`, `norvexa-scan`).
- Options: `"meaning": false` (skip the back-translation check), `"term_packs": ["eu-ctr"]`, `"stream": true`.
- Limits: 2 to 8 documents, 30,000 characters each and 120,000 in total, at most 2 scans or PDFs of 8 MB and 4 pages.
- The JSON response has `status` (drift, check or consistent), `documents` (with the sections found), `pairs` (what was
compared with what), `flags` (each with `pair`, `section`, `kind`, `severity`, `what`, `reference` and `other` quotes
with `doc`, `text`, `start`, `end` and, for scans, `cite`; `judgement` with `label`, `reason` and the deviation-log line
it relied on), `counts`, `warnings`, `record` and `receipt_events`.
- With `Accept: text/event-stream` (or `"stream": true`): `reading`/`read` per scan, `ready`, a `receipt` per model call,
a `pair` event per document pair with its flags, `result`, `budget` and `done`. A seven-document set takes a few minutes.
- `POST /record/verify` (no token) `{"record": {...}}` re-checks the signed record.
- `GET /label/info` (sources, regional conventions, QRD headings, limits), `GET /label/samples` and `GET /attest/signing-key` need no token.
## Example (Python, `pip install httpx`)
```python
import httpx, os, pathlib
API = "https://api.decosa.ai"
H = {"Authorization": f"Bearer {os.environ['DECOSA_API_KEY']}"}
docs = [("ccds", "en", "ccds.txt"), ("us_pi", "en", "us_pi.txt"), ("eu_smpc", "en", "smpc_en.txt"), ("eu_smpc", "de", "smpc_de.txt")]
body = {"documents": [{"role": r, "lang": l, "text": pathlib.Path(f).read_text()} for r, l, f in docs],
"declared_deviations": [], "synthetic": True}
r = httpx.post(f"{API}/label/check", json=body, headers=H, timeout=900)
r.raise_for_status()
js = r.json()
for f in js["flags"]:
if f["judgement"]["label"] != "likely_deliberate":
print(f["judgement"]["label"], f["section"], f["what"], "|", (f["reference"] or {}).get("text"), "|", (f["other"] or {}).get("text"))
```
Run it yourself (containers)
On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.
# Decosa label consistency check: run it yourself (containers)
You are setting up Decosa's label consistency check on this machine, so unapproved labelling never leaves it. It aligns
the sections of a CCDS, a US PI, an EU SmPC and leaflet in several languages and carton text or a scan, flags every
difference with a quote from each document, checks translations (numbers, negations, meaning by back-translation) and
SmPC headings against the QRD template, triages each flag from the regional conventions and your deviation log, and
signs a record. Nothing is sent to Decosa's hosted API. It is a labelling QC aid: regulatory affairs decides.
Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.
Ask me before any command that needs sudo, and show me the command first.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/label-consistency-check.zip (8 KB, 7 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py label-consistency-check` (the api image carries the same bundle under /app/rehearsal/label-consistency-check/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py label-consistency-check --bundle label-consistency-check.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the set is reported as drifted", "the US PI against the CCDS has at least two likely errors in its warnings (the 6-month interval and the missing warning)", "the dropped German negation is flagged in the translation"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
instructions for this distribution (docs.docker.com/engine/install). Install the NVIDIA container toolkit and check
`docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
`mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
Read it. Keep the `llm` service (Qwen3.8-27B on vLLM) and the `api` service. For the `api` service set
`DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b` and
`DECOSA_LABELCHECK_SYNTHETIC_ONLY=0` (so this box accepts real labelling), and bind every port to 127.0.0.1. Never set
the gateway route on a box that holds real labelling.
3. Translations: add the language pack's translation model (Hy-MT2-7B on vLLM, about 18 GB) and set
`DECOSA_LANG_MT_URL`, or send `"meaning": false` (numbers, negations and QRD headings are still checked in code).
4. Scanned cartons (optional): add the document reader (`services/docreader` plus PaddleOCR-VL-1.6 on vLLM, about 6 GB)
and set `DECOSA_DOCREADER_URL`. Without it, send carton text.
5. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check.
6. Check: `curl -fsS http://127.0.0.1:<PORT>/label/info` shows `synthetic_only: false`; `GET /attest/signing-key` shows
this box's public key.
7. Smoke test: get a token with `POST /demo/session {"vertical":"label-consistency-check"}` and send
`{"sample_id": "norvexa-quick"}` to `POST /label/check`. Expect `status: drift`, a likely error in warnings (liver tests
every 6 months against 3 in the CCDS) and no likely error in indications (that difference is in the sample's deviation
log). Send the record to `POST /record/verify`: `ok` must be true.
8. Report back: the public key, the flag counts, and how long the run took.
Off by default. Joining as a provider serves other people's requests on this GPU; never do it on a box that holds a
company's labelling. If I ask for it later, follow the Provide page instead of improvising.
Run it on your own hardwareWhat it needs, and the prompt that sets it up
Run it on your own GPU
Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.
Hardware check
Check your own hardware- CPU only, 64 GB RAMDoesn't fit
Qwen3.8-27B (NVFP4, vision tower on) needs a GPU.
- GeForce RTX 4090lite tierRuns with a smaller tier
The standard tier does not fit: Needs about 44 GB of GPU memory at the smallest settings; 24 GB available. The lite tier fits with changes.
- GeForce RTX 5090lite tierRuns with a smaller tier
The standard tier does not fit: Needs about 52 GB of GPU memory at the smallest settings; 32 GB available. The lite tier fits with changes.
- 2x GeForce RTX 5090standard tierRuns
The standard tier fits with changes: Qwen3.8-27B (NVFP4, vision tower on): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
- L40Slite tierRuns with a smaller tier
The standard tier does not fit: Needs about 57.6 GB of GPU memory at the smallest settings; 48 GB available. The lite tier fits with changes.
- H100 80 GB (SXM)best tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4, vision tower on) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU. The best tier fits too.
- RTX PRO 6000 Blackwell 96 GBbest tierRuns
The standard tier fits (81.6 of 96 GB). The best tier fits too.
- 2x RTX PRO 6000 Blackwell 96 GBbest tierRuns
The standard tier fits (81.6 of 192 GB). The best tier fits too.
- Apple M3 Ultra (Mac Studio), 96 GBlite tierRuns with a smaller tier
The standard tier can't be checked: Hy-MT2-7B has no mapped Apple Silicon build The lite tier fits with changes.
- Apple M5 Max, 64 GBlite tierRuns with a smaller tier
The standard tier can't be checked: Hy-MT2-7B has no mapped Apple Silicon build The lite tier fits with changes.
Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.
On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.
- 1
Check the GPU, Docker and the NVIDIA Container Toolkit
The driver must see the GPU, and Docker must be able to pass it into a container.
nvidia-smi docker compose version docker run --rm --gpus all ubuntu nvidia-smi
- 2
Fetch the compose file
One file describes the API and the language model as services.
mkdir -p ~/decosa && cd ~/decosa curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml - 3
Pull and start
The first start downloads pinned model weights, tens of gigabytes.
docker compose pull docker compose up -d
- 4
Check health
Wait until the API reports ok with the language model loaded. Then point your app at the local base URL.
curl -fsS http://localhost:<PORT>/healthz # {"ok": true, "llm": true, ...} curl -fsS -X POST http://localhost:<PORT>/demo/session \ -H 'Content-Type: application/json' -d '{"vertical":"label-consistency-check"}'
Set up with a coding agent, rehearse on mock data, then go private
- Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
- Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's
expected.json. Every check must print PASS. - Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
docker compose exec api python scripts/rehearse.py label-consistency-check
Download the mock-data bundle (8 KB, 7 checks)expected.json
Synthetic labels for Norvexa (tavorexin), an invented medicine. The US PI's liver-test interval says every 6 months where the CCDS says 3, the US PI has lost the depression and suicidal ideation warning, and the German SmPC has dropped the negation in 'must not be initiated in patients with an active serious infection'. The US-only indication and pregnancy wording are in the deviation log. The check must flag the drifts as likely errors, keep the logged differences out of the errors, leave the German QRD headings alone, attach a receipt to every model call and sign a record that verifies.
What the rehearsal checks
- the set is reported as drifted
- the US PI against the CCDS has at least two likely errors in its warnings (the 6-month interval and the missing warning)
- the dropped German negation is flagged in the translation
- the logged US-only indication difference is not called an error
- the German SmPC's QRD headings are all right
- every model call has a signed receipt
- the signed record verifies
Licence: Label texts written for this bundle (CC0); the medicine, company and every number are invented. Part of decosa-api, AGPL-3.0-or-later.
Prompt for your coding agent
# Decosa label consistency check: run it yourself (containers)
You are setting up Decosa's label consistency check on this machine, so unapproved labelling never leaves it. It aligns
the sections of a CCDS, a US PI, an EU SmPC and leaflet in several languages and carton text or a scan, flags every
difference with a quote from each document, checks translations (numbers, negations, meaning by back-translation) and
SmPC headings against the QRD template, triages each flag from the regional conventions and your deviation log, and
signs a record. Nothing is sent to Decosa's hosted API. It is a labelling QC aid: regulatory affairs decides.
Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.
Ask me before any command that needs sudo, and show me the command first.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/label-consistency-check.zip (8 KB, 7 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py label-consistency-check` (the api image carries the same bundle under /app/rehearsal/label-consistency-check/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py label-consistency-check --bundle label-consistency-check.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the set is reported as drifted", "the US PI against the CCDS has at least two likely errors in its warnings (the 6-month interval and the missing warning)", "the dropped German negation is flagged in the translation"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
instructions for this distribution (docs.docker.com/engine/install). Install the NVIDIA container toolkit and check
`docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
`mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
Read it. Keep the `llm` service (Qwen3.8-27B on vLLM) and the `api` service. For the `api` service set
`DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b` and
`DECOSA_LABELCHECK_SYNTHETIC_ONLY=0` (so this box accepts real labelling), and bind every port to 127.0.0.1. Never set
the gateway route on a box that holds real labelling.
3. Translations: add the language pack's translation model (Hy-MT2-7B on vLLM, about 18 GB) and set
`DECOSA_LANG_MT_URL`, or send `"meaning": false` (numbers, negations and QRD headings are still checked in code).
4. Scanned cartons (optional): add the document reader (`services/docreader` plus PaddleOCR-VL-1.6 on vLLM, about 6 GB)
and set `DECOSA_DOCREADER_URL`. Without it, send carton text.
5. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check.
6. Check: `curl -fsS http://127.0.0.1:<PORT>/label/info` shows `synthetic_only: false`; `GET /attest/signing-key` shows
this box's public key.
7. Smoke test: get a token with `POST /demo/session {"vertical":"label-consistency-check"}` and send
`{"sample_id": "norvexa-quick"}` to `POST /label/check`. Expect `status: drift`, a likely error in warnings (liver tests
every 6 months against 3 in the CCDS) and no likely error in indications (that difference is in the sample's deviation
log). Send the record to `POST /record/verify`: `ok` must be true.
8. Report back: the public key, the flag counts, and how long the run took.
Off by default. Joining as a provider serves other people's requests on this GPU; never do it on a box that holds a
company's labelling. If I ask for it later, follow the Provide page instead of improvising.
Help me customise for my hardware
Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.
GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page
Runs with a smaller tierLabel consistency across PI, SmPC, CCDS and carton on GeForce RTX 5090: use the Lite · text documents, meaning check off tier
The standard tier does not fit: Needs about 52 GB of GPU memory at the smallest settings; 32 GB available. The lite tier fits with changes.
What this tool's stack says about this hardware:
- 1x RTX 5090 32 GB (does not fit): Estimate: Qwen3.8-27B NVFP4 with a small KV cache fits, the translation model and reader do not; run the lite tier (text documents, meaning check off) or put them on a second card.
Lite · text documents, meaning check off: what changesuses estimates
- Qwen3.8-27B (NVFP4, vision tower on): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
Memory per component
- Section splitter: decosa-api label check (decosa_api/verticals/labelcheck), on the language-pack block (decosa_api/lang), the document reader (decosa_api/docreader) and the signed record (07). CPU. Runs on CPU (vram_gb 0 in stack.json).
- One call per section pair: Qwen3.8-27B (NVFP4, vision tower on). ~57.6 GB (at least ~28 GB), weights 21.4 GB (from stack.json). Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)
Expected speed
Not measured.
Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.
Setup prompt for this hardware
The self-host prompt for Label consistency across PI, SmPC, CCDS and carton, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.
# Set up Label consistency across PI, SmPC, CCDS and carton on my hardware Fetch https://decosa.ai/prompts/label-consistency-check-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied. ## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=label-consistency-check) Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4). Quality tier: Lite · text documents, meaning check off (lite). Fit check: runs with changes, about 28 GB of 32 GB used; some memory numbers are estimates, not measurements. First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything. Use these components (the setup below describes the standard tier; change it to match): - Section splitter: decosa-api label check (decosa_api/verticals/labelcheck), on the language-pack block (decosa_api/lang), the document reader (decosa_api/docreader) and the signed record (07), CPU - One call per section pair: Qwen3.8-27B (NVFP4, vision tower on) (nvidia/Qwen3.8-27B-NVFP4), 57.6 GB. Change: Qwen3.8-27B (NVFP4, vision tower on): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions. GPU placement (set each service's device and its vLLM --gpu-memory-utilization to about the share shown): - GPU 0: Qwen3.8-27B (NVFP4, vision tower on) ~28 GB (88%); about 4 GB left During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed. The stack's own component list and compose layout: https://decosa.ai/prompts/label-consistency-check-assemble.md
The proof
How we tested itEval results and end-to-end checks, hosted and self-hosted, with dates
Verified end to end
Hosted: verified 28 Sep 2026 · measured 28 Sep 2026: · p50 2.3 s · p95 24 s (6 runs) · ~$0.001 per run · 3 receipts
Loading the nightly status…
Self-host: verified 27 Sep 2026 · fresh clone into a clean directory, api image from docker/api/Dockerfile, compose with a named volume, direct route to the local Qwen3.8-27B, language pack and document reader, local signing; torn down after
Measured cost to run: about $0.015 per label set (hosted, 28 Sep 2026, partly estimated). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.
The rehearsal bundle passed 7/7 in 13.8 s; every receipt signed. The language pack and reader were the running services, not built from the compose file here.
Known limits (5)
- Hosted numbers are from 6 production smoke runs after the merge: 3 on 27-28 Sep under load (18.2, 21.0, 24.3 s) and 3 on a quiet gateway on 28 Sep (2.1, 2.3, 2.3 s). Cost is the median at list price, model calls included (range $0.0011 to $0.0011).
- Measured on synthetic label sets written by the same author as the prompts; not on a real company's labels or against a labelling reviewer's findings.
- Full seven-document sets take minutes on the shared gateway (415 s measured); the Watch replay shows a recorded real run.
- QRD headings are checked in English, German and French only; other EU languages get the number lock and meaning check.
- Interactions, pregnancy sections, pharmacology and artwork layout are not compared.
How it's builtThe steps, the models and what each one checks
Run it yourself, on request
- The same open models and app, on 1x RTX PRO 6000 (96 GB): Qwen3.8-27B NVFP4 plus Hy-MT2-7B (about 18 GB) and the document reader (about 6 GB); sections, number checks and the record run on CPU.
- Data never leaves your machines, and there are no Decosa charges.
- One prompt for Claude Code or Codex assembles the whole stack.
- Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Get an API key
- Call the label consistency across pi, smpc, ccds and carton API from your own code in minutes.
- Every model answer carries a signed receipt.
- Synthetic, public or test data only: real confidential data belongs on your own hardware.
A CCDS, a US PI, an EU SmPC and leaflet in several languages and the carton in; every difference in the safety sections out, quoted from each document and triaged as likely error, likely deliberate or unclear.
For regulatory affairs and labelling teams. It aligns the sections of each label document (indications, dosing, contraindications, warnings, adverse reactions, storage, strengths), compares the US PI and the English SmPC with the company core data sheet, the leaflet and carton with the SmPC, and each translation with its English text. Every difference comes back with a quote from each document: numbers with units are compared in code, translations are checked for numbers, negations and changed meaning (back-translation plus a model judgment), and SmPC headings against the EMA QRD template. Each flag is triaged from the regional conventions and your deviation log. A labelling QC aid with a signed record; regulatory affairs decides.
- Deployment
- Hosted or self-host
- Regulatory
- Checked 27 Sep 2026. US: 21 CFR 201.57(c) sets the Full Prescribing Information sections (1 Indications and usage, 2 Dosage and administration, 3 Dosage forms and strengths, 4 Contraindications, 5 Warnings and precautions, 6 Adverse reactions ... 16 How supplied/storage and handling; eCFR, current), and 21 CFR 314.70(c)(6)(iii) lets a holder add or strengthen a contraindication, warning, precaution or adverse reaction, or a dosage instruction for safe use, on a changes-being-effected (CBE-0) basis. EU: Directive 2001/83/EC Article 11 sets the SmPC headings in order (4.1 therapeutic indications ... 4.3 contra-indications, 4.4 special warnings ..., 4.8 undesirable effects, 6.4 special precautions for storage) and Article 63(1) requires labelling and leaflet particulars in the official language(s) of the Member State; read on legislation.gov.uk (the EU text as it stood on 31 Dec 2020), so later amendments and the EU pharmaceutical reform are unverified here. The EMA QRD product-information template v10.4 (02/2024; a draft v11 was consulted on until 31 Aug 2025) gives the headings and standard statements in each language; the check uses its English, German and French versions. ICH E2C(R2) (Step 4, 17 Dec 2012) describes the Company Core Data Sheet and its core safety information (CCSI), which is why the CCDS is the reference. The check lists and triages differences; it does not decide whether a label is approvable, whether a deviation is justified or which regulatory route a change needs. Not legal or regulatory advice. Model licences: Apache-2.0 (Qwen3.8-27B, Hy-MT2-7B, PaddleOCR-VL-1.6, Docling Heron layout); MIT (Docling).
Text description
Label documents (CCDS, US PI, SmPC in several languages, leaflet, carton text or scan, and a deviation log) go to decosa-api. The document reader turns scans into text with page and box. Code splits each document into sections and plans the pairs. Qwen3.8-27B compares each section pair and returns quotes from both documents; code locates the quotes and compares numbers with units. The language pack checks translations: numbers and negations in code, meaning by Hy-MT2-7B back-translation and a Qwen3.8 judgment. The QRD term pack checks SmPC headings. Each flag is triaged as likely error, unclear or likely deliberate, and a signed record lists the document hashes, flags and receipt ids; regulatory affairs decides.
At a glance
- Data retention
- Nothing stored: the documents live in memory for the request. The signed record holds document hashes, sections, each flag's place, kind and triage and the receipt ids, never the label text; logs carry counts and timings only.
- What leaves the box
- Self-hosted on the direct route: nothing. The model, the translation model, the document reader and the checks run on the same machine. Hosted: model calls go through the Decosa API, and only invented or public labels are accepted.
- What it compares
- Indications, dosing, contraindications, warnings, adverse reactions, storage and strengths: US PI and English SmPC against the CCDS, leaflet and carton against the SmPC, each language against its English text, SmPC headings against the QRD template (English, German, French).
- What it is not
- A regulatory decision, an approval of a deviation or an artwork proofreader. Interactions, pregnancy sections, pharmacology and layout are not compared. Regulatory affairs decides each flag.
Pick the tier for the quality you need
Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.
Lite
text documents, meaning check off
Text documents only, sent with "meaning": false: the model compares sections and triages; translations get the code checks only (numbers, units, negations, QRD headings), no back-translation. Only the language model needs a GPU.
- Models
- decosa-api label check (decosa_api/verticals/labelcheck), on the language-pack block (decosa_api/lang), the document reader (decosa_api/docreader) and the signed record (07)
- Qwen3.8-27B (NVFP4, vision tower on)
- Hardware
- 1x GPU for Qwen3.8-27B
- Quality evidence
- Planted drifts caughtnot measured yetthe eval ran with the meaning check on; 2 of 9 translation drifts on the test split were found only by the meaning check
- Latency
- not measured
- Verification
- Proof: strongSelf-host onlyEvery model call is receipted.
- In the hosted demo
Standard
Qwen3.8-27B, Hy-MT2-7B and the document reader (hosted demo)
What the hosted demo runs: Qwen3.8-27B compares and triages, Hy-MT2-7B back-translates for the meaning check, the document reader reads scanned cartons, code checks numbers, headings and quotes and signs.
- Models
- decosa-api label check (decosa_api/verticals/labelcheck), on the language-pack block (decosa_api/lang), the document reader (decosa_api/docreader) and the signed record (07)
- Qwen3.8-27B (NVFP4, vision tower on)
- Hy-MT2-7B
- Document reader block: Docling 2.130 (Heron layout) + PaddleOCR-VL-1.6 (0.9B)
- Hardware
- 1x RTX PRO 6000 96 GB (measured on shared cards)
- Quality evidence
- Planted drifts caught, held-out synthetic set (numbers, missing warnings and contraindications, storage, translations)31 / 31, all triaged likely error; 30 / 31 quoted to a phrase or sentencedecosa-api docs/evals/label-consistency-check.md, test split (Pelmora, run once)
- Same drift with an explaining deviation-log entry, triaged likely deliberate6 / 6decosa-api docs/evals/label-consistency-check.md, test split
- False flags on the clean held-out set3 of 24 flags (all from the translation number lock; 0 after a post-test fix, re-scored on the same model outputs)decosa-api docs/evals/label-consistency-check.md
- Deliberate-or-error triage against a blind frontier judge (Claude Opus 5.5, same prompt), 72 casesQwen3.8 72 / 72, Opus 71 / 72; 35 / 35 each on the test casesdecosa-api docs/evals/label-consistency-check.md
- Full seven-document set, hosted gateway route415 s under shared load, 43 Qwen3.8 calls, $0.0147 at list pricedecosa-api docs/evals/label-consistency-check.md
- Latency
- measured on the shared gateway: several minutes for the full document set, a couple of minutes for a pair or a scanned carton, about a minute for the smoke check; seconds self-hosted (direct route)
- Verification
- Proof: partialQwen3.8 calls: gateway receipts. Back-translation and reader calls: model-call attestations and a signed reader receipt from decosa-api, not countersigned by the gateway yet.
Best
the 30B translation model (not measured)
Hy-MT2-30B-A3B-FP8 for back-translation; not run by us.
- Models
- decosa-api label check (decosa_api/verticals/labelcheck), on the language-pack block (decosa_api/lang), the document reader (decosa_api/docreader) and the signed record (07)
- Qwen3.8-27B (NVFP4, vision tower on)
- Document reader block: Docling 2.130 (Heron layout) + PaddleOCR-VL-1.6 (0.9B)
- Hy-MT2-30B-A3B-FP8
- Hardware
- 1x RTX PRO 6000 96 GB (estimate: 20 GB Qwen3.8 NVFP4 weights plus 31 GB MT)
- Quality evidence
- Planted translation drifts caughtnot measured yetnot run
- Latency
- not measured
- Verification
- Proof: partialSelf-host only
Every model in the stack
| Model | Tiers | Params · VRAM | Verification | Details |
|---|---|---|---|---|
Section splitter (QRD numbers, 21 CFR 201.57 numbers, heading words), pair plan, quote location at character offsets, number-with-unit comparison, QRD heading check, triage rules for translations, signed record (no model; CPU)decosa-api label check (decosa_api/verticals/labelcheck), on the language-pack block (decosa_api/lang), the document reader (decosa_api/docreader) and the signed record (07) 0 GBProof: partial | LiteStandardBest | 0 GB | Proof: partial | |
| ||||
One call per section pair: the differences as JSON with an exact quote from each document; one triage call per document pair; the meaning judgment per translated paragraph; a full-page read of a scanned carton (image input)Qwen3.8-27B (NVFP4, vision tower on)nvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab) 27B · 20 GBProof: strong | LiteStandardBest | 27B · 20 GB | Proof: strong | |
| ||||
Back-translation of each translated paragraph into English for the meaning check (the language-pack block): German, French and the other languages it servesHy-MT2-7Btencent/Hy-MT2-7B on Hugging Face (opens in a new tab) 7.5B · 18 GBProof: partial | Standard | 7.5B · 18 GB | Proof: partial | |
| ||||
Scanned or PDF cartons and labels: finds and orders the regions of each page (Docling Heron layout) and reads them (PaddleOCR-VL-1.6), with a page and box per lineDocument reader block: Docling 2.130 (Heron layout) + PaddleOCR-VL-1.6 (0.9B)PaddlePaddle/PaddleOCR-VL-1.6 on Hugging Face (opens in a new tab) 0.9B · 6 GBProof: partialIn the hosted demo | StandardBest | 0.9B · 6 GB | Proof: partialIn the hosted demo | |
| ||||
Back-translation, larger modelHy-MT2-30B-A3B-FP8tencent/Hy-MT2-30B-A3B-FP8 on Hugging Face (opens in a new tab) 30B (3B active) · about 31 GB (estimate)No proof yet | Best | 30B (3B active) · about 31 GB (estimate) | No proof yet | |
| ||||
Tools, services and hardware
Tools
- EMA QRD product-information template v10.4 (EN, DE, FR) (opens in a new tab)EMA legal notice: reproduction for commercial and non-commercial purposes with acknowledgement
The SmPC headings and standard statements in the eu-qrd-human term pack (read 27 Sep 2026).
- EMA EPAR product information (Otezla, DE and FR) (opens in a new tab)EMA legal notice: reproduction with acknowledgement
Only the four adverse-reaction frequency category names (sehr häufig / très fréquent ...), so the meaning check accepts them.
- 21 CFR 201.57 and 314.70 (eCFR) (opens in a new tab)US federal regulation (public domain)
The US PI section numbers the splitter reads; the CBE-0 route cited in the notes.
- Synthetic label sets of invented medicinesAGPL-3.0-or-later (decosa-api, decosa_api/verticals/labelcheck/data)
The samples and the eval: Norvexa (dev) and Pelmora (test), each with a CCDS, US PI, SmPC in English, German and French, an English leaflet and a carton, with 31 planted drifts and 6 declared twins per medicine.
Services
- decosa-api:8445
${DECOSA_REGISTRY}/decosa-api:<tag>GET /label/info, /label/samples; POST /label/check (SSE or JSON); POST /record/verify. Keeps no label text.
- decosa-llm:8000
${DECOSA_REGISTRY}/decosa-llm:0.1.0vLLM OpenAI endpoint for Qwen3.8-27B with image input.
- decosa-lang-mt:8491
vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1Hy-MT2-7B, the language pack's translation model, for back-translation.
- decosa-docreader:8497
built from services/docreader (no published image yet), with PaddleOCR-VL-1.6 on vLLMReads scanned or PDF cartons into lines with page and box.
Hardware
- 1x RTX PRO 6000 Blackwell 96 GB Fits
Measured on our server with the models on shared cards: Qwen3.8-27B NVFP4 (about 20 GB of weights), Hy-MT2-7B (about 18 GB) and the document reader (about 6 GB).
- 1x RTX 5090 32 GB Does not fit
Estimate: Qwen3.8-27B NVFP4 with a small KV cache fits, the translation model and reader do not; run the lite tier (text documents, meaning check off) or put them on a second card.
- CPU only Does not fit
The model needs a GPU. Sections, number checks and the record run on CPU.
Latency per lane
- the seven-document Norvexa set (CCDS, US PI, SmPC EN/DE/FR, leaflet, carton), hosted gateway route414.7 s
Measuredmeasured on our server 2026-09-27 (pre-release server, gateway shared with other workloads, load average about 65)
- CCDS against US PI (norvexa-quick), hosted gateway route147.7 s
Measuredmeasured on our server 2026-09-27 (pre-release server)
- the four-document rehearsal set, self-hosted direct route to the local Qwen3.813.8 s
Measuredmeasured on our server 2026-09-27 (fresh clone, compose)
Notes
- Numbers with units, QRD headings, quotes and translation figures and negations are checked in code; the model finds and describes section differences and triages them against the regional conventions and your deviation log.
- Held out (synthetic): 31 of 31 planted drifts caught and triaged as likely errors, 6 of 6 logged twins triaged as deliberate, 3 false flags in 24 on the clean set (all fixed after the test).
- A labelling QC aid: it does not approve a deviation, decide whether a label is approvable or choose a regulatory route; regulatory affairs decides each flag.
Run this exact stack on your machine
Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.
# Assemble the Decosa label consistency check on this machine
You are setting up a self-hosted labelling QC aid on this Linux machine, for a regulatory affairs or labelling team. It
takes the label documents for one medicine (the company core data sheet (CCDS), the US Prescribing Information, the EU
SmPC and package leaflet in several languages, and carton text or a carton scan) and returns:
- the sections of each document aligned (indications, dosing, contraindications, warnings, adverse reactions, storage,
strengths);
- every difference with a quote from each document at a character offset: a dose, interval, strength or temperature that
disagrees, a missing warning or contraindication, a translation that dropped a negation or changed a number, an SmPC
heading that is not the QRD template's;
- a triage per difference (likely error, likely deliberate, unclear) using the regional conventions it knows and our
deviation log;
- a signed, hash-chained record of the document hashes, flags, triage and every model receipt id.
Work step by step, show me each command before running anything that needs sudo, and stop if a check fails.
**Before anything else, remind me:**
- Unapproved labelling and variation drafts are confidential. On this box nothing leaves the machine: the language model,
the translation model, the document reader and the checks all run here. Never point it at a hosted gateway while it
holds real labelling.
- It is a QC aid, not a regulatory decision. The triage is a model judgment; regulatory affairs decides what is deliberate
and records approved deviations. It compares seven sections only (not interactions, pregnancy, pharmacology or artwork
layout), and QRD headings in English, German and French only.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/label-consistency-check.zip (8 KB, 7 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py label-consistency-check` (the api image carries the same bundle under /app/rehearsal/label-consistency-check/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py label-consistency-check --bundle label-consistency-check.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the set is reported as drifted", "the US PI against the CCDS has at least two likely errors in its warnings (the 6-month interval and the missing warning)", "the dropped German negation is flagged in the translation"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## What you are building
| service | image | model | port |
|---|---|---|---|
| `llm` | `${DECOSA_REGISTRY}/decosa-llm:0.1.0` (vLLM 0.29.0) | `nvidia/Qwen3.8-27B-NVFP4` @ `482ca0f3832238542f8f5295dde86b5f22711d80`, Apache-2.0 | internal 8000 |
| `mt` | `vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1` | `tencent/Hy-MT2-7B` @ `9b0eb4e8f001def3e5ff6469a0ac96fdb39ec223`, Apache-2.0 | internal 8000 |
| `parser` | same vLLM image | `PaddlePaddle/PaddleOCR-VL-1.6` @ `c5630abae1d940eafe0697512a0325494b02ab42`, Apache-2.0 | internal 8000 |
| `docreader` | built from `decosa-api/services/docreader` | Docling 2.130 (MIT) with `docling-project/docling-layout-heron` @ `8f39ad3c0b4c58e9c2d2c84a38465abf757272d8` (Apache-2.0) | internal 8497 |
| `api` | `${DECOSA_REGISTRY}/decosa-api:0.1.0` (no GPU) | none | `127.0.0.1:8445` |
Only text documents (no scans or PDFs)? Leave out `parser` and `docreader` and their `depends_on` entry, and drop
`DECOSA_DOCREADER_URL`. Only English documents? You can leave out `mt` too and send `"meaning": false`; numbers and
negations in translations are still checked in code.
## 1. Check the GPU, driver and Docker
1. `nvidia-smi`: one NVIDIA GPU with at least 64 GB (Qwen3.8-27B NVFP4 about 20 GB plus KV cache, Hy-MT2-7B about 18 GB
at `--gpu-memory-utilization 0.18`, the document reader about 6 GB), driver 580 or newer. NVFP4 needs Blackwell; on
Hopper set `LLM_MODEL=Qwen/Qwen3.8-27B-FP8` and `LLM_REVISION=main` (not measured). Smaller cards: tell me.
2. `docker --version`, `docker compose version`, `docker run --rm --gpus all ubuntu nvidia-smi`. If Docker or the NVIDIA
Container Toolkit is missing, install them from the official Docker and NVIDIA repositories after asking me.
3. About 80 GB of free disk.
## 2. Images and source
The images are **on request** while self-host is in early access: ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry (set it as `DECOSA_REGISTRY`), pull access and the compose file. try `docker pull ${DECOSA_REGISTRY}/decosa-{llm,api}:0.1.0`. If a pull
fails, clone `decosa-api` (access required) into `~/decosa-label/decosa-api`, check out the newest tag that contains
`decosa_api/verticals/labelcheck/` (`main` until one does), and build `docker build -f docker/api/Dockerfile -t
${DECOSA_REGISTRY}/decosa-api:0.1.0 .`. The `docreader` service is always built from the clone.
## 3. The compose file
`~/decosa-label/.env`:
```bash
DECOSA_TAG=0.1.0
DECOSA_GPU=0
LLM_MODEL=nvidia/Qwen3.8-27B-NVFP4
LLM_REVISION=482ca0f3832238542f8f5295dde86b5f22711d80
LLM_GPU_UTIL=0.55
DECOSA_SIGNER_NAME="<who signs these records, e.g. Example Pharma labelling QC>"
```
`~/decosa-label/docker-compose.yml`:
```yaml
name: decosa-label
x-health: &health { interval: 15s, timeout: 5s, retries: 5 }
x-gpu: &gpu { resources: { reservations: { devices: [ { driver: nvidia, device_ids: ["${DECOSA_GPU:-0}"], capabilities: [gpu] } ] } } }
services:
llm:
image: ${DECOSA_REGISTRY}/decosa-llm:${DECOSA_TAG}
deploy: *gpu
ipc: host
restart: unless-stopped
volumes: [hf-cache:/root/.cache/huggingface]
command: ["${LLM_MODEL}", "--revision", "${LLM_REVISION}", "--served-model-name", "qwen3.8-27b", "--max-model-len", "32768",
"--gpu-memory-utilization", "${LLM_GPU_UTIL}", "--limit-mm-per-prompt", '{"image":4}', "--max-num-seqs", "16",
"--seed", "0", "--host", "0.0.0.0", "--port", "8000"]
healthcheck: { <<: *health, test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8000/health', timeout=4)"], start_period: 900s }
mt:
image: vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1
deploy: *gpu
ipc: host
restart: unless-stopped
volumes: [hf-cache:/root/.cache/huggingface]
# vLLM 0.29 serves Hy-MT2 through its Transformers backend; CUDA-graph capture fails there, hence --enforce-eager
command: ["tencent/Hy-MT2-7B", "--revision", "9b0eb4e8f001def3e5ff6469a0ac96fdb39ec223", "--served-model-name", "hy-mt2-7b",
"--max-model-len", "8192", "--gpu-memory-utilization", "0.18", "--max-num-seqs", "16", "--enforce-eager",
"--generation-config", "vllm", "--host", "0.0.0.0", "--port", "8000"]
healthcheck: { <<: *health, test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8000/health', timeout=4)"], start_period: 600s }
parser:
image: vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1
deploy: *gpu
ipc: host
restart: unless-stopped
volumes: [hf-cache:/root/.cache/huggingface]
command: ["PaddlePaddle/PaddleOCR-VL-1.6", "--revision", "c5630abae1d940eafe0697512a0325494b02ab42", "--served-model-name", "PaddlePaddle/PaddleOCR-VL-1.6",
"--trust-remote-code", "--max-model-len", "8192", "--gpu-memory-utilization", "0.04", "--max-num-seqs", "16",
"--no-enable-prefix-caching", "--mm-processor-cache-gb", "0", "--generation-config", "vllm", "--host", "0.0.0.0", "--port", "8000"]
healthcheck: { <<: *health, test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8000/health', timeout=4)"], start_period: 600s }
docreader:
build: { context: ./decosa-api/services/docreader }
deploy: *gpu
restart: unless-stopped
depends_on: { parser: { condition: service_healthy } }
environment: { DOCREADER_PARSER: paddle-openai, DOCREADER_PARSER_URL: "http://parser:8000/v1", DOCREADER_PARSER_MODEL: PaddlePaddle/PaddleOCR-VL-1.6 }
volumes: [hf-cache:/models]
healthcheck: { <<: *health, test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8497/health', timeout=4)"], start_period: 300s }
api:
image: ${DECOSA_REGISTRY}/decosa-api:${DECOSA_TAG}
restart: unless-stopped
depends_on: { llm: { condition: service_healthy }, mt: { condition: service_healthy }, docreader: { condition: service_healthy } }
environment:
DECOSA_LLM_ROUTE: direct # local model only; receipts are signed by this box's key ("attested")
DECOSA_LLM_URL: http://llm:8000/v1
DECOSA_LLM_MODEL: qwen3.8-27b
DECOSA_LANG_MT_URL: http://mt:8000/v1
DECOSA_LANG_MT_MODEL: hy-mt2-7b
DECOSA_DOCREADER_URL: http://docreader:8497
DECOSA_LOCAL_SIGNING: "on" # Ed25519 key created at /data/attest/ on first start
DECOSA_SIGNER_NAME: ${DECOSA_SIGNER_NAME}
DECOSA_LABELCHECK_SYNTHETIC_ONLY: "0" # this box accepts real labelling
DECOSA_SESSIONS_PER_IP_HOUR: "1000"
DECOSA_BUDGET_LLM_TOKENS: "400000" # per session; a seven-document set uses about 15,000 generated tokens
DECOSA_SESSION_TTL_S: "28800"
ports: ["127.0.0.1:8445:8445"]
volumes: [decosa-data:/data]
healthcheck: { <<: *health, test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8445/label/info', timeout=4)"], start_period: 20s }
volumes: { hf-cache: {}, decosa-data: {} }
```
Use the named volumes exactly as written: the api image runs as an unprivileged user, and a root-owned host bind mount
makes it fail on `/data/keys.sqlite`. Run `docker compose up -d --build` and poll `docker compose ps` until every service
is healthy (the LLM takes 5-10 minutes the first time). Back up the signing key with
`docker compose cp api:/data/attest ./attest-backup`; keep it private and never print it.
## 4. Smoke test on the bundled invented medicine
```bash
API=localhost:8445
curl -s $API/label/info | jq '{synthetic_only, qrd: .qrd_headings["4.3"]}'
TOKEN=$(curl -s $API/demo/session -H 'content-type: application/json' -d '{"vertical":"label-consistency-check"}' | jq -r .token)
curl -s $API/label/check -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' -d '{"sample_id":"norvexa-quick"}' > /tmp/q.json
jq -r '.status, (.flags[] | "\(.judgement.label)\t\(.section)\t\(.what)")' /tmp/q.json
jq '{record}' /tmp/q.json | curl -s $API/record/verify -H 'content-type: application/json' -d @- | jq '{ok, summary}'
curl -s $API/label/check -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' -d '{"sample_id":"norvexa-drift"}' | jq '.status, .counts'
curl -s $API/label/check -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' -d '{"sample_id":"norvexa-scan"}' | jq '.status, .warnings, [.flags[].other.text]'
```
Pass if: `synthetic_only` is `false` and the 4.3 heading is `Gegenanzeigen` in German; the quick sample is `drift` with a
likely error in warnings (the liver-test interval, 6 months against 3) and none in indications (that difference is in
the sample's deviation log); the record verifies; the full set is `drift` with at least five likely errors; and the scan
flags the carton's 25°C against the SmPC's 30°C. Every receipt should be `attested`. The quick sample takes about 30 s
and the full set a few minutes; tell me what you measure.
## 5. Check our own labels
```bash
python3 - <<'PY' > req.json
import json, pathlib
docs = [("ccds", "en", "ccds.txt"), ("us_pi", "en", "uspi.txt"), ("eu_smpc", "en", "smpc_en.txt"), ("eu_smpc", "de", "smpc_de.txt")]
print(json.dumps({"documents": [{"role": r, "lang": l, "text": pathlib.Path(f).read_text()} for r, l, f in docs],
"declared_deviations": [l.strip() for l in pathlib.Path("deviations.txt").read_text().splitlines() if l.strip()]}))
PY
curl -s $API/label/check -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' -d @req.json > qc.json
jq .record qc.json > label-check-record.json
```
Keep section headings in the text (4.3 Contraindications, 5 WARNINGS AND PRECAUTIONS): sections are found by the QRD and
21 CFR 201.57 numbers, and by heading words for a CCDS. Export Word or PDF labels to text first; a carton can go as a scan
(`{"role": "carton", "file_b64": "...", "media_type": "image/png"}`, up to 2 files per run). Limits: 8 documents, 30,000
characters each. Put each known, approved regional difference on one line of the deviation log.
## 6. Point tools at the local API
- Base URL `http://localhost:8445`; for a web app set `NEXT_PUBLIC_DECOSA_API=http://localhost:8445`.
- `POST /label/check` returns JSON, or Server-Sent Events with `Accept: text/event-stream`. `GET /label/info` lists the
sources, the regional conventions, the QRD headings and what it does not check.
- Keep the record with the change-control file; anyone can re-check it with `POST /record/verify` against the key at
`GET /attest/signing-key`. Keep the API on 127.0.0.1; for other users, put a TLS reverse proxy with authentication in front.
## 7. Keep the direct route
`DECOSA_LLM_ROUTE=gateway` would send label text to the hosted Decosa API. Leave it off for real labelling.Rules and regulations it checks againstDated, linked to the primary source; not legal advice
Regulation watch
Loading the watch status…
5 laws, rules and guidance pages cited; 5 watched nightly at the primary source. A change marks this page for a human re-check; nothing is edited automatically. What we cite and how it is watched
Technical detailsModels, where it runs, labels
In short
Last reviewed
- What it is
- A CCDS, a US PI, an EU SmPC and leaflet in several languages and the carton in; every difference in the safety sections out, quoted from each document and triaged as likely error, likely deliberate or unclear.
- Who it's for
- Regulatory affairs and labelling teams at pharma and biotech companies, and labelling service providers.
- Where it runs
- Self-host for unapproved labelling; the hosted demo takes invented or public labels only
- Key numbers
- 31 / 31 Planted drifts caught (test split, n = 31)
- 30 / 31 Quotes tight to the drift (test split, n = 31)
- 6 / 6 Logged twins triaged deliberate (test split, n = 6)
- 2.3 s Median end-to-end run, hosted (QA sweep 2026-09-28)
- Models
- Qwen3.8-27B (section comparison with quotes, triage, meaning judgment) · Hy-MT2-7B (back-translation, language-pack block) · Docling layout + PaddleOCR-VL (scanned cartons, document reader)
- Where
- Self-host for unapproved labelling; the hosted demo takes invented or public labels only
- Checks
- Receipt per model call; every quote found in its document at a character offset; numbers with units compared in code; signed hash-chained record of the document hashes, flags and triage
- Industry
- Healthcare · Compliance and trust
- Output
- Structured data · Signed record or verdict
- Data
- Confidential business data
- Hardware
- 1× 96 GB GPU
- Licence
- Permissive (Apache-2.0, MIT)
- Runs in
- Decosa hosted · Self-host
- Built from
- Language pack · Document reader · Signed record
Questions people ask
Do our unapproved labels leave our network?
Not when you self-host: the model, the translation model, the document reader and the checks run on your own GPU box. The hosted demo accepts invented or public labels only and keeps no label text.
How does it tell a deliberate difference from an error?
From the regional conventions it knows (US °F and controlled room temperature, US lab units, QRD standard statements, lay leaflet wording) and your deviation log. On the held-out synthetic set, 6 of 6 logged differences were triaged deliberate and 31 of 31 planted drifts as likely errors. Without a log entry, a missing warning is called a likely error.
Which differences does it look for?
Doses, strengths, units, intervals and percentages that disagree; missing or weakened warnings and contraindications; storage conditions; and translation drift: changed numbers, dropped negations, missing paragraphs and changed meaning, plus SmPC headings against the QRD template (English, German, French).
Does it read scanned cartons?
Yes, through the document reader block: a scanned or PDF carton is read into lines with page and box, then compared with the SmPC.
Does it approve labels or decide the regulatory route?
No. It is a labelling QC aid: it lists and quotes differences and gives a first triage. Regulatory affairs decides each flag and whether a change needs a variation or supplement.
Ask a question or leave feedbackWe read every message and publish useful answers
Ask about Label consistency across PI, SmPC, CCDS and carton
We read every message. Questions, comments and our answers show here once we have reviewed and approved them.
Loading questions…