Pre-check promo claims for MLR
An annotated MLR pre-review packet: each claim with the reference passage behind it, plus off-label, comparative, fair-balance and disease-claim problems.
Built on: Grounding, Signed record
Loading the tool…
Use it your way
Use it from your codeThe hosted API with your key, and prompts to paste into a coding agent
Get an API key
- Call the promotional-claims pre-check API from your own code in minutes.
- Every model answer carries a signed receipt.
- Nothing to install; we run the models.
Run it yourself, on request
- The same open models and app, on 1× RTX PRO 6000 (96 GB) for the model; extraction, rules and the packet run on CPU.
- Data never leaves your machines, and there are no Decosa charges.
- One prompt for Claude Code or Codex assembles the whole stack.
- Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Build with it
Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.
- Base URL
- https://api.decosa.ai
- Auth
Authorization: Bearer $DECOSA_API_KEY(or a demo session token)- Tool id
- promo-claims-check
Use the hosted API
# Decosa promotional-claims pre-check: use the hosted API
You are wiring Decosa's promotional-claims pre-check into this project. It takes a promotional piece for a prescription
drug, a medical device or a dietary supplement (the text) and its reference pack (the label or prescribing information,
the cited studies, spec sheets, and optionally the approved claims matrix). It checks every sentence: is it supported by
the references (with the span), is it off-label, is a comparison backed by head-to-head data, is a superlative backed,
and for supplements, is it a disease claim. It then checks the whole piece for fair balance, the boxed warning and the
DSHEA disclaimer, and returns an MLR pre-review packet in Markdown plus a packet signed by the server. Every model call
has its own signed receipt. Use only what is listed below. If you need something else, stop and ask me.
- Base URL: `https://api.decosa.ai`
- Health check: `GET https://api.decosa.ai/healthz`.
- This is a pre-check that speeds up MLR review. It does not approve anything and does not replace the committee.
Say so wherever you show results.
- The hosted API is for public labels and synthetic or already published copy. Unpublished study data and pre-launch
pieces belong on a self-hosted box.
## Auth: API key (or a demo session)
1. Preferred: an API key (`dk_…`) from "Get an API key" on the tool page. Keep it in an environment variable,
`DECOSA_API_KEY`, never in code. Send `Authorization: Bearer $DECOSA_API_KEY`.
2. Without a key: `POST https://api.decosa.ai/demo/session` with `{"vertical": "promo-claims-check"}` returns `{"token", "expires_at", "budget"}`.
a limited number of sessions per network per hour (the current limits are in `demo_sessions` of GET /healthz); each session has a token allowance (its `budget`). Over a limit: HTTP 429 with `Retry-After`.
3. A check needs about 330 generated tokens per claim sentence plus 250 (402 otherwise), so a demo session covers two
or three pieces of 15 claims. One check at a time per demo token.
## Endpoints
- `POST /promo/check` (token). Body:
```json
{"piece": {"text": "...", "title": "..."},
"product": {"type": "drug|device|supplement", "name": "...", "indication": "optional; read from the label otherwise"},
"references": [{"title": "...", "kind": "label|study|other", "text": "..."}],
"claims_matrix": [{"id": "M1", "claim": "..."}]}
```
- Limits: piece up to 8,000 characters and 40 claim sentences; up to 10 references and 200,000 characters in total;
up to 60 matrix entries. Send text, not URLs or PDFs: extract the text yourself.
- The approved indication comes from `product.indication`, else from the INDICATIONS AND USAGE (or Indications for
Use) section of a `kind: "label"` reference. With neither, the off-label check does not run (`no_indication`).
- JSON by default: `{status: "issues"|"checks"|"clean", counts, indication, indication_source, claims: [{i, text, role,
support, confidence, evidence: [{span, source, start, end, role, quote}], findings, support_reason, review_reason,
comparative, superlative, off_label, claim_kind, matrix, h2h, receipt_ids}], piece: {findings, fair_balance},
receipts, packet, packet_md, budget, note}`.
- With `Accept: text/event-stream` (or `"stream": true`) it streams `ready`, `receipt` events and one `claim` event
per sentence (not in text order), then `piece`, `packet`, `budget`, `done`.
- `POST /promo/verify` (no token) `{"packet": {...}, "text"?: "...", "references"?: [...], "packet_md"?: "..."}` →
`{valid_signature, signed_by_this_server, status, piece_matches?, references_match?, packet_md_matches?}`.
- `GET /promo/info` (the finding codes, severities and limits), `GET /promo/samples`, `GET /attest/signing-key`.
Finding codes. Issues (likely violations): `unsupported`, `contradicted`, `off_label`, `comparative_unbacked`,
`superlative_unbacked`, `disease_claim`, `not_checked`, `fair_balance_missing`, `fair_balance_weak`,
`boxed_warning_missing`, `dshea_disclaimer_missing`. Checks (for the reviewer): `overstated`, `not_in_matrix`,
`pi_reference_missing`, `no_indication`.
## Example: pre-check a piece before it goes into the review queue (Python, `pip install httpx`)
```python
import httpx, os, pathlib
API = "https://api.decosa.ai"
H = {"Authorization": f"Bearer {os.environ['DECOSA_API_KEY']}"}
body = {"piece": {"text": piece_text, "title": "Spring HCP email"},
"product": {"type": "drug", "name": "Metformin hydrochloride tablets"},
"references": [{"title": "Prescribing information", "kind": "label", "text": label_text}],
"claims_matrix": [{"id": "M1", "claim": "..."}]}
r = httpx.post(f"{API}/promo/check", json=body, headers=H, timeout=600)
r.raise_for_status()
res = r.json()
pathlib.Path(f"{res['packet']['id']}.md").write_text(res["packet_md"]) # attach this to the review job
pathlib.Path(f"{res['packet']['id']}.json").write_text(r.text) # keep the signed packet with the piece
for c in res["claims"]:
if c["findings"]:
print(c["findings"], c["text"], "--", c["support_reason"])
```
## Verify a packet yourself (`pip install cryptography`)
```python
import hashlib, json, urllib.request
from cryptography.hazmat.primitives.asymmetric.ed25519 import Ed25519PublicKey
pub = json.load(urllib.request.urlopen("https://api.decosa.ai/attest/signing-key"))["pubkey"]
pk = res["packet"]
assert pk["signer"] == pub
body = {k: v for k, v in pk.items() if k != "sig"}
msg = pk["v"].encode() + b"\n" + json.dumps(body, sort_keys=True, separators=(",", ":"), ensure_ascii=False).encode()
Ed25519PublicKey.from_public_bytes(bytes.fromhex(pub)).verify(bytes.fromhex(pk["sig"]), msg)
assert hashlib.sha256(piece_text.encode()).hexdigest() == pk["piece"]["sha256"]
assert hashlib.sha256(res["packet_md"].encode()).hexdigest() == pk["packet_md_sha256"]
```
The signed packet holds hashes, offsets, finding codes and receipt ids, never the copy. It embeds the signed grounding
report for the support verdicts. Each receipt id resolves at `GET https://api.decosa.ai/receipts/{id}`.
## Honest limits
- Text only: type size, placement and contrast (visual prominence) are not assessed. Fair balance is measured as risk
text against benefit text plus whether the boxed warning and contraindications are conveyed.
- Findings come from a language model. On our planted-violation eval it caught most planted problems and raised some
issues on compliant sentences; see the numbers on the Stack tab. A person reviews every finding.
- Not legal or regulatory advice.
Run it yourself (containers)
On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.
# Decosa promotional-claims pre-check: run it yourself (containers)
You are setting up the Decosa promotional-claims pre-check on this machine, so pre-launch pieces and unpublished study
data never leave it. It checks every claim of a promotional piece against its reference pack and returns a signed MLR
pre-review packet. Nothing is sent to Decosa's hosted API.
Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.
Ask me before any command that needs sudo, and show me the command first.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/promo-claims-check.zip (5 KB, 10 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py promo-claims-check` (the api image carries the same bundle under /app/rehearsal/promo-claims-check/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py promo-claims-check --bundle promo-claims-check.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the piece comes back with issues", "at least two disease claims are flagged", "the planted 'cure insomnia' sentence is flagged as a disease claim"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
instructions for this distribution (docs.docker.com/engine/install). Install the NVIDIA container toolkit and check
`docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
`mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
Read it. Keep the `llm` service (Qwen3.8-27B on vLLM, with prefix caching on) and the `api` service. For the `api`
service set `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b` and bind
every port to 127.0.0.1. Never set the gateway route on this box: it would send the copy to the Decosa API.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check (the first start
downloads about 20 GB of weights).
4. Check: `curl -fsS http://127.0.0.1:<PORT>/promo/info` lists the finding codes and limits; `GET /attest/signing-key`
shows this box's public key. Show me the key: it is what reviewers pin to verify my packets.
5. Smoke test: get a token with `POST /demo/session {"vertical":"promo-claims-check"}`, fetch `GET /promo/samples`, and
send the `metformin-planted` sample (piece, product, references, claims_matrix) to `POST /promo/check`. Expect
`status: "issues"`, `off_label` on the weight-loss and heart-protection sentences, a comparison finding on "better
than any other diabetes pill", and `fair_balance_missing` at piece level. Then send `metformin-compliant`: expect no
issue-level findings. Then `POST /promo/verify` with the packet and the packet_md: `valid_signature`,
`signed_by_this_server` and `packet_md_matches` must all be true.
6. Report back: the public key and key id, both statuses, the findings and how long each check took.
Off by default. Joining as a provider serves other people's requests on this GPU; never do it on a box that holds
pre-launch pieces or unpublished data. If I ask for it later, follow the Provide page instead of improvising.
No NVIDIA GPU? This tool also runs entirely on an Apple Silicon Mac (MLX, 32 GB of unified memory or more): use https://decosa.ai/prompts/promo-claims-check-mac.md instead.
Run it on your own hardwareWhat it needs, and the prompt that sets it up
Run it on your own GPU
Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.
Hardware check
Check your own hardware- CPU only, 64 GB RAMDoesn't fit
Qwen3.8-27B (NVFP4) needs a GPU.
- GeForce RTX 4090standard tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)
- GeForce RTX 5090standard tierRuns
The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
- 2x GeForce RTX 5090standard tierRuns
The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size).
- L40Sstandard tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.
- H100 80 GB (SXM)standard tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.
- RTX PRO 6000 Blackwell 96 GBstandard tierRuns
The standard tier fits (57.6 of 96 GB).
- 2x RTX PRO 6000 Blackwell 96 GBstandard tierRuns
The standard tier fits (57.6 of 192 GB).
- Apple M3 Ultra (Mac Studio), 96 GBstandard tierRuns
The standard tier fits (32 of 96 GB).
- Apple M5 Max, 64 GBstandard tierRuns
The standard tier fits (32 of 64 GB).
Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.
On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.
- 1
Check the GPU, Docker and the NVIDIA Container Toolkit
The driver must see the GPU, and Docker must be able to pass it into a container.
nvidia-smi docker compose version docker run --rm --gpus all ubuntu nvidia-smi
- 2
Fetch the compose file
One file describes the API and the language model as services.
mkdir -p ~/decosa && cd ~/decosa curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml - 3
Pull and start
The first start downloads pinned model weights, tens of gigabytes.
docker compose pull docker compose up -d
- 4
Check health
Wait until the API reports ok with the language model loaded. Then point your app at the local base URL.
curl -fsS http://localhost:<PORT>/healthz # {"ok": true, "llm": true, ...} curl -fsS -X POST http://localhost:<PORT>/demo/session \ -H 'Content-Type: application/json' -d '{"vertical":"promo-claims-check"}'
Set up with a coding agent, rehearse on mock data, then go private
- Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
- Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's
expected.json. Every check must print PASS. - Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
docker compose exec api python scripts/rehearse.py promo-claims-check
Download the mock-data bundle (5 KB, 10 checks)expected.json
A product page for a fictional magnesium supplement, checked against the public NIH fact sheet and a synthetic spec. It makes three disease claims (insomnia, blood pressure, migraines) and has no DSHEA disclaimer, so the pre-check must flag them, and the signed MLR packet must verify and catch a changed status.
What the rehearsal checks
- the piece comes back with issues
- at least two disease claims are flagged
- the planted 'cure insomnia' sentence is flagged as a disease claim
- the missing DSHEA disclaimer is flagged
- the signed packet verifies
- the packet matches the MLR Markdown packet
- the packet matches the piece
- the packet matches the references
- a packet with its status changed to clear no longer verifies
- every model call has a signed receipt
Licence: Synthetic promotional copy for the fictional brand Fernhill, written for Decosa. Reference: NIH Office of Dietary Supplements magnesium fact sheet (US government work, public domain) plus a synthetic product spec. Part of decosa-api, AGPL-3.0-or-later.
Prompt for your coding agent
# Decosa promotional-claims pre-check: run it yourself (containers)
You are setting up the Decosa promotional-claims pre-check on this machine, so pre-launch pieces and unpublished study
data never leave it. It checks every claim of a promotional piece against its reference pack and returns a signed MLR
pre-review packet. Nothing is sent to Decosa's hosted API.
Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.
Ask me before any command that needs sudo, and show me the command first.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/promo-claims-check.zip (5 KB, 10 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py promo-claims-check` (the api image carries the same bundle under /app/rehearsal/promo-claims-check/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py promo-claims-check --bundle promo-claims-check.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the piece comes back with issues", "at least two disease claims are flagged", "the planted 'cure insomnia' sentence is flagged as a disease claim"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
instructions for this distribution (docs.docker.com/engine/install). Install the NVIDIA container toolkit and check
`docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
`mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
Read it. Keep the `llm` service (Qwen3.8-27B on vLLM, with prefix caching on) and the `api` service. For the `api`
service set `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b` and bind
every port to 127.0.0.1. Never set the gateway route on this box: it would send the copy to the Decosa API.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check (the first start
downloads about 20 GB of weights).
4. Check: `curl -fsS http://127.0.0.1:<PORT>/promo/info` lists the finding codes and limits; `GET /attest/signing-key`
shows this box's public key. Show me the key: it is what reviewers pin to verify my packets.
5. Smoke test: get a token with `POST /demo/session {"vertical":"promo-claims-check"}`, fetch `GET /promo/samples`, and
send the `metformin-planted` sample (piece, product, references, claims_matrix) to `POST /promo/check`. Expect
`status: "issues"`, `off_label` on the weight-loss and heart-protection sentences, a comparison finding on "better
than any other diabetes pill", and `fair_balance_missing` at piece level. Then send `metformin-compliant`: expect no
issue-level findings. Then `POST /promo/verify` with the packet and the packet_md: `valid_signature`,
`signed_by_this_server` and `packet_md_matches` must all be true.
6. Report back: the public key and key id, both statuses, the findings and how long each check took.
Off by default. Joining as a provider serves other people's requests on this GPU; never do it on a box that holds
pre-launch pieces or unpublished data. If I ask for it later, follow the Provide page instead of improvising.
No NVIDIA GPU? This tool also runs entirely on an Apple Silicon Mac (MLX, 32 GB of unified memory or more): use https://decosa.ai/prompts/promo-claims-check-mac.md instead.
Help me customise for my hardware
Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.
GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page
RunsPromotional-claims pre-check on GeForce RTX 5090: use the Standard · one GPU for the model (hosted demo) tier
The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
What this tool's stack says about this hardware:
- 1x RTX 5090 32 GB (fits): Estimate: Qwen3.8-27B NVFP4 needs about 20 GB of weights plus KV cache; not run for this use case. Reference packs of 20k+ characters make long prompts, so keep prefix caching on.
Standard · one GPU for the model (hosted demo): what changesuses estimates
- Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
Memory per component
- Pre-check: decosa-api promo module (decosa_api/verticals/promo) with the grounding module (decosa_api/verticals/grounding). CPU. Runs on CPU (vram_gb 0 in stack.json).
- Grounding judge, claim reviewer, head-to-head...: Qwen3.8-27B (NVFP4). ~57.6 GB (at least ~28 GB), weights 21.4 GB (from stack.json). Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)
Expected speed
Not measured.
Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.
Setup prompt for this hardware
The self-host prompt for Promotional-claims pre-check, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.
# Set up Promotional-claims pre-check on my hardware Fetch https://decosa.ai/prompts/promo-claims-check-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied. ## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=promo-claims-check) Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4). Quality tier: Standard · one GPU for the model (hosted demo) (standard). Fit check: runs with changes, about 28 GB of 32 GB used; some memory numbers are estimates, not measurements. First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything. Use these components (the setup below describes the standard tier; change it to match): - Pre-check: decosa-api promo module (decosa_api/verticals/promo) with the grounding module (decosa_api/verticals/grounding), CPU - Grounding judge, claim reviewer, head-to-head...: Qwen3.8-27B (NVFP4) (nvidia/Qwen3.8-27B-NVFP4), 57.6 GB. Change: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions. GPU placement (set each service's device and its vLLM --gpu-memory-utilization to about the share shown): - GPU 0: Qwen3.8-27B (NVFP4) ~28 GB (88%); about 4 GB left During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed. The stack's own component list and compose layout: https://decosa.ai/prompts/promo-claims-check-assemble.md
Or on a Mac Studio
No NVIDIA GPU needed: every model this tool uses runs natively on Apple Silicon through MLX. Any M-series Mac with 32 GB of unified memory or more. Measured speeds and what runs where
From a checkout of decosa-api, one command sets up the models and the API: scripts/mac/setup.sh
Mac prompt for your coding agent
# Decosa Promotional-claims pre-check: run it on this Mac (Apple Silicon, no NVIDIA GPU) You are setting up the Decosa Promotional-claims pre-check on this Mac, natively on Apple Silicon. The models run on the Mac's GPU through MLX and decosa-api runs from a git checkout with `uv`. Docker is not used for the models, because Docker on macOS cannot reach the GPU. Nothing is sent to Decosa's hosted API. Every model this tool needs runs on the Mac. It needs 32 GB of unified memory or more. Ask me before any command that needs sudo or installs software with Homebrew, and show me the command first. Never stop or kill a process this setup did not start; if a port is taken, pick another one. ## Step 0: set up with a coding agent, rehearse on mock data, then go private This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works. Work in this order: 1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to "test with something realistic". 2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool, https://decosa.ai/samples/promo-claims-check.zip (5 KB, 10 checks, synthetic or openly licensed: see `licence` in expected.json), show me what is in it, and run the rehearsal against the local API: `.venv/bin/python scripts/rehearse.py promo-claims-check` in the decosa-api checkout (the key comes from ~/.decosa-mac/api.key). It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the piece comes back with issues", "at least two disease claims are flagged", "the planted 'cure insomnia' sentence is flagged as a disease claim"). Show me the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json` to make a check pass. 3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this machine. For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent can read. Switch to your own data only after the rehearsal has passed and the agent's work is done. ## What runs where | Part | On an NVIDIA GPU | On this Mac | Status | |---|---|---|---| | Pre-check: sentences, label sections, rules, findings, the packet (no model; CPU) | Python on CPU | The same Python module, run with uv | Runs, measured | | Grounding judge, claim reviewer, head-to-head and fair-balance checks | NVFP4 on vLLM 0.29 (Blackwell) | MLX 4-bit (EigenLabs/Qwen3.8-27B-4bit) on mlx_lm.server 0.31.3; oMLX 0.6.1 with MTP as an option | Runs, measured | ## Steps 1. Check the machine: `uname -m` must print `arm64` (an M-series chip; Intel Macs cannot run MLX), and `sysctl -n hw.memsize` should be at least 32 GB for this tool. Check about 30 GB of free disk with `df -h ~`. Show me the chip (`sysctl -n machdep.cpu.brand_string`) and the memory. 2. Tools: `uv --version`. If it is missing, ask me, then `brew install uv`. 3. Code: `git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host> ~/decosa-api` (access required) and `cd ~/decosa-api`. Check that `scripts/mac/setup.sh` exists; if it does not, the checkout is too old: stop and tell me. 4. Start everything with one command: `scripts/mac/setup.sh`. It creates `.venv` (decosa-api) and `.venv-mac` (MLX, mlx-lm, mlx-audio), downloads the weights with the Hugging Face CLI (about 16 GB for the language model), starts the model servers and decosa-api on 127.0.0.1, and mints a local API key into `~/.decosa-mac/api.key` (mode 0600). The first run takes a while because of the downloads; later runs reuse them. If a download fails with 401 or 403, ask me for a Hugging Face token and set `HF_TOKEN`. 5. Check health: `scripts/mac/setup.sh status` shows each server, and `curl -fsS http://127.0.0.1:8445/healthz` must report `"llm": true`. `curl -fsS http://127.0.0.1:8445/attest/signing-key` shows this Mac's public key: show it to me, because it is what others pin to check the receipts and records this Mac signs. 6. Smoke test: `.venv/bin/python scripts/mac/bench_usecases.py promo-claims-check`. It runs the tool's own sample end to end against the local API with the local key and prints `ok`, the wall time, the model calls and the receipts. `ok=True` is the pass condition. If it fails, read `~/.decosa-mac/logs/*.log` and tell me what you found. 7. Point the app at it: the API is `http://127.0.0.1:8445` with `Authorization: Bearer $(cat ~/.decosa-mac/api.key)`, the same routes as the hosted API. To stop everything: `scripts/mac/setup.sh stop`. 8. Report back: the chip and memory, the public key, the smoke-test result and its time, and the output of `scripts/mac/setup.sh status`. ## Good to know - Receipts: every model call is signed with this Mac's own Ed25519 key and names the exact MLX weights (`qwen3.8-27b-mlx-4bit` with a hash of the downloaded files). There is no gateway countersignature on a self-hosted Mac. - The weights are a 4-bit MLX build of the same open models, not the NVFP4 build the hosted route and the published evals use. Expect small differences in wording and scores. - Faster drafting: `scripts/mac/setup.sh stop && scripts/mac/setup.sh --engine omlx` serves the model with oMLX and multi-token prediction (about 2x faster for a single long answer, no faster for many parallel calls; typed judgments then use sampling because oMLX returns no log-probabilities). - Measured speeds for a Mac Studio M3 Ultra and the memory each tool needs: https://decosa.ai/mac. Full details: `docs/self-host-mac.md` in the checkout.
The proof
How we tested itEval results and end-to-end checks, hosted and self-hosted, with dates
Verified end to end
Hosted: verified 25 Sep 2026 · measured 25 Sep 2026: · p50 3.9 s · ~$0.022 per run · 14 receipts
Loading the nightly status…
Self-host: verified 25 Sep 2026 · fresh clone, compose up, sample against local model servers
Measured cost to run: about $0.022 per piece (hosted, 25 Sep 2026). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.
Method: a fresh clone of decosa-api main, the api image built from it, the compose file from this prompt, then the prompt's smoke steps and the nightly smoke module, against the already-running local Qwen3.8-27B vLLM. Verified on 25 Sep 2026: images build, services start, sample passes end to end against local model servers equivalent to the documented ones; model-server startup itself not re-verified. Planted metformin piece: status issues with every expected finding (contradicted HbA1c, off-label weight and heart claims, comparative, fair balance, boxed warning); the compliant piece came back clean; the packet verifies and a changed finding fails.
Known limits (3)
- Speed depends on load: a 7-14 sentence piece took 4-10 s on a quiet GPU and 30-35 s while the shared GPU was busy (25 Sep 2026).
- Text only: type size, placement and contrast are not assessed. Tables flattened to text can be misread.
- The eval's planted problems are blatant ones; subtle violations are not measured yet. A person reviews every finding.
How it's builtThe steps, the models and what each one checks
Get an API key
- Call the promotional-claims pre-check API from your own code in minutes.
- Every model answer carries a signed receipt.
- Nothing to install; we run the models.
Run it yourself, on request
- The same open models and app, on 1× RTX PRO 6000 (96 GB) for the model; extraction, rules and the packet run on CPU.
- Data never leaves your machines, and there are no Decosa charges.
- One prompt for Claude Code or Codex assembles the whole stack.
- Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Every claim in a drug, device or supplement piece checked against its references before MLR review, with the span behind it and what is wrong.
Send a promotional piece and its reference pack: the label or prescribing information, the cited studies, spec sheets and, optionally, the approved claims matrix. Each sentence is checked against the references by the grounding module, which returns the supporting span. A second receipted call types the claim: off-label against the approved indication, comparative or superlative, and disease or structure/function for supplements. The whole piece is then checked for fair balance, the boxed warning and the DSHEA disclaimer. The output is an annotated MLR pre-review packet, plus a signed record of hashes and findings. It is a pre-check that shortens review, not a replacement for it.
- Deployment
- Hosted or self-host
- Regulatory
- Checked 25 Sep 2026. This is a pre-check that helps a medical-legal-regulatory (MLR) committee; it approves nothing, and the company and its reviewers stay responsible for every piece. What it checks against: FDA's prescription drug advertising rules (21 CFR 202.1: not false or misleading, fair balance of benefit and risk, only uses in the approved labeling), which FDA's Office of Prescription Drug Promotion enforces; FDA's 2018 guidance on communications consistent with the FDA-required labeling; the DSHEA rules for supplement claims (21 CFR 101.93: structure/function claims need the disclaimer in 101.93(c), disease claims are not allowed); and the FTC's Health Products Compliance Guidance (December 2022), which expects competent and reliable scientific evidence that fits each claim. It reads text only, so the visual prominence of risk information is not assessed, and a claim the references support can still be misleading in context. Findings come from a language model and can be wrong in both directions (see the eval). Unpublished study data and pre-launch pieces should stay on your own hardware: self-host. The hosted demo keeps nothing (piece and references stay in memory for the request; the packet holds hashes; logs carry counts) and should be used with public labels and synthetic or published copy only. Model licence: Apache-2.0 (Qwen3.8-27B). Not legal or regulatory advice.
Text description
A promotional piece, its reference pack (label, studies, spec sheets) and an optional claims matrix go to the pre-check, which splits the piece into sentences, numbers the reference spans and reads the approved indication, boxed warning and contraindications from the label. Qwen3.8-27B makes a grounding call and a review call per claim, a head-to-head call for comparative drug and device claims and a fair-balance call per piece. Rules check the DSHEA disclaimer and the pointer to full prescribing information. Outputs: per-claim findings with the reference span, piece-level findings, and an MLR pre-review packet in Markdown with a signed JSON record. On the hosted route every model call gets a signed receipt that our gateway countersigns. In self-host mode everything runs on your machine.
At a glance
- Data retention
- Nothing stored: the piece and references live in memory for the request. The signed packet holds hashes, offsets, finding codes and receipt ids, never the copy.
- What leaves the box
- Hosted: every model call goes through our gateway to the GPU serving Qwen3.8-27B, and its receipt (hashes, token counts, no text) is kept by the gateway and this API. Self-hosted on the direct route: nothing leaves the box.
- Input formats
- Text only: piece up to 8,000 characters and 40 claim sentences; up to 10 references (label, studies, other) and 200,000 characters; up to 60 claims-matrix entries. Extract PDF text yourself.
- Typical run
- The metformin sample: a dozen or so model calls and a few cents at the gateway list price; the supplement sample about the same. Each run shows its own measured cost.
Pick the tier for the quality you need
Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.
- In the hosted demo
Standard
one GPU for the model (hosted demo)
Qwen3.8-27B does the grounding, the claim review and the head-to-head and fair-balance checks, each in its own receipted call. This is what the hosted API runs.
- Models
- decosa-api promo module (decosa_api/verticals/promo) with the grounding module (decosa_api/verticals/grounding)
- Qwen3.8-27B (NVFP4)
- Hardware
- 1x RTX PRO 6000 96 GB (measured) or 1x RTX 5090 32 GB (estimate)
- Quality evidence
- Planted problems caught, held-out test (unsupported, off-label, comparative, disease)20 / 20 (9/9, 3/3, 3/3, 5/5), same in a repeat rundocs/evals/promo-claims-check.md: 8 synthetic pieces on 3 public labels and 3 NIH fact sheets
- Piece-level problems caught, test (fair balance, DSHEA disclaimer)2 / 2docs/evals/promo-claims-check.md
- False positives, test: unplanted sentences with an issue / with any finding1 / 63 (1.6%) / 3-4 of 63docs/evals/promo-claims-check.md, two runs
- Dev split (the only split used for a change)11 / 11 caught; 2 / 31 unplanted with an issuedocs/evals/promo-claims-check.md
- Latency
- measured: under a minute to over a minute per piece through the shared gateway.
- Verification
- Proof: strongEvery model call is a separate gateway call with a gateway-signed receipt; the signed packet lists them all.
Also runs on
- Layout-aware checkQwen3.8-27B vision inputnot builtRead the rendered page as well as its text with the 27B's own vision input, so fair balance can account for type size and placement. Not built. Hardware: 1x RTX PRO 6000 96 GB (estimate).
We host these ourselves when needed: small models get more of our own compute unless we detect a shortage, so they need no community providers.
Every model in the stack
| Model | Tiers | Params · VRAM | Verification | Details |
|---|---|---|---|---|
Pre-check: sentences, label sections, rules, findings, the packet (no model; CPU)decosa-api promo module (decosa_api/verticals/promo) with the grounding module (decosa_api/verticals/grounding) 0 GBProof: partial | Standard | 0 GB | Proof: partial | |
| ||||
Grounding judge, claim reviewer, head-to-head and fair-balance checksQwen3.8-27B (NVFP4)nvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab) 27.8B · 20 GBProof: strongIn the hosted demo | Standard | 27.8B · 20 GB | Proof: strongIn the hosted demo | |
| ||||
Layout reader for visual prominence (alternate)Qwen3.8-27B vision inputnvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab) 27.8B · 20 GBNo proof yetSelf-host only | Alternate | 27.8B · 20 GB | No proof yetSelf-host only | |
| ||||
Tools, services and hardware
Tools
- openFDA drug label API / DailyMed (opens in a new tab)Public domain (US government data; openFDA terms)
The metformin, atorvastatin and lisinopril labels used in the demo and eval (section excerpts; set ids in scripts/promo_cases.py).
- NIH Office of Dietary Supplements fact sheets (opens in a new tab)Public domain (US government work)
Magnesium, vitamin D and omega-3 health-professional fact sheets as supplement references.
- scripts/promo_eval.py and docs/evals/promo-claims-check.mdApache-2.0
12 synthetic pieces with planted violations (dev and test split), the scoring and the results.
- POST /promo/verifyApache-2.0
Checks a packet's signature against this server's key and, if you send them, the hashes of the piece, the references and the Markdown packet. The console also checks the signature in your browser with WebCrypto.
Services
- decosa-api:8445
${DECOSA_REGISTRY}/decosa-api:<tag>GET /promo/info, /promo/samples; POST /promo/check (SSE or JSON), /promo/verify. Keeps no text.
- vLLM (model):8114
vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1Qwen3.8-27B NVFP4 behind our gateway (hosted) or called directly (self-host).
Hardware
- 1x RTX PRO 6000 Blackwell 96 GB Fits
Measured on our server: the hosted demo and the eval ran on this card, shared with other services.
- 1x RTX 5090 32 GB Fits
Estimate: Qwen3.8-27B NVFP4 needs about 20 GB of weights plus KV cache; not run for this tool. Reference packs of 20k+ characters make long prompts, so keep prefix caching on.
Latency per lane
- one piece of 7-14 sentences, hosted gateway route35.0 s
Measuredmeasured on our server 2026-09-25: 25-77 s per piece over 12 eval pieces and 8 demo recordings, shared gateway and GPU
- model calls per piecen/a
Measuredmeasured 2026-09-25: about 20 calls and 1.6k generated tokens per piece (156 calls, 13.0k tokens for 8 test pieces)
Notes
- The planted problems in the eval are blatant ones, written by the same agent that built the checker; 20 of 20 caught means obvious problems are caught, not that subtle ones are. The next eval should use real OPDP untitled and warning letters, which are public.
- The one test false positive, in both test runs: a trial result in a population the label covers was called off-label. On copy written to comply, 1 of 63 sentences drew an issue and 3-4 drew a check.
- Tables flattened to text are a weak spot: while recording the demo, the compliant metformin piece once had a head-to-head sentence marked contradicted because the judge misread the label's flattened table (1 of 4 runs of that piece).
- Rx drugs and devices: a comparative claim needs head-to-head evidence in the references (a separate receipted call checks the cited passages). Supplements: a comparison the references state outright is backed, since the FTC asks for evidence that fits the claim rather than a head-to-head trial.
- The signed packet holds hashes, offsets, finding codes and receipt ids; the Markdown packet that quotes the copy is returned to you and its hash is in the signature.
Run this exact stack on your machine
Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.
# Assemble the Decosa promotional-claims pre-check on this machine
You are setting up a pre-check for promotional pieces (prescription drugs, medical devices, dietary supplements). It
takes the piece's text and its reference pack (label or prescribing information, cited studies, spec sheets, and
optionally the approved claims matrix), checks every claim against the references, and returns an annotated MLR
pre-review packet plus a JSON packet signed by this box's own key. Work step by step, show me each command before you
run anything with `sudo`, and stop to ask if a check fails.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/promo-claims-check.zip (5 KB, 10 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py promo-claims-check` (the api image carries the same bundle under /app/rehearsal/promo-claims-check/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py promo-claims-check --bundle promo-claims-check.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the piece comes back with issues", "at least two disease claims are flagged", "the planted 'cure insomnia' sentence is flagged as a disease claim"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## 0. Ground rules and licences
- Model: Qwen3.8-27B (Apache-2.0), used for three jobs: the grounding judge (is the claim in the references), the claim
review (role, off-label, comparative, superlative, disease vs structure/function, claims-matrix match), and the
head-to-head and fair-balance checks. The pre-check itself is decosa-api (AGPL-3.0-or-later) and needs no GPU of its own.
- Pieces and references stay on this machine. Bind every port to 127.0.0.1. The service keeps no text: nothing is
written to disk and logs carry counts only. Keep it that way; do not add request logging.
- Be honest about what it is: a pre-check that speeds up MLR review, not a replacement for it. It approves nothing,
reads text only (no visual prominence), and its findings come from a language model that can be wrong both ways.
## 1. Check the machine
1. `nvidia-smi`: one GPU with at least 32 GB (Qwen3.8-27B NVFP4 needs about 20 GB of weights plus KV cache; an RTX PRO
6000 96 GB is what we measured on, an RTX 5090 32 GB should fit but we have not run this tool on one). Driver
570 or newer. Blackwell cards run NVFP4; on older cards use `Qwen/Qwen3.8-27B-FP8`.
2. `docker --version` and `docker compose version`. If Docker or the NVIDIA container toolkit is missing, install them
from the official Docker and NVIDIA repositories after asking me, then run
`docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
3. Disk: about 30 GB free.
## 2. Images and weights
- `${DECOSA_REGISTRY}/decosa-api:<tag>` (**publishing soon**). If the pull fails, build from source:
`git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host>` (access required), check out the newest release tag that contains
`decosa_api/verticals/promo/` (`main` until one does: v0.1.0 predates it) and `decosa_api/verticals/grounding/`, and build `docker/api/Dockerfile`.
- `vllm/vllm-openai:v0.29.0` for the model; weights `nvidia/Qwen3.8-27B-NVFP4` (revision
`482ca0f3832238542f8f5295dde86b5f22711d80`), or `Qwen/Qwen3.8-27B-FP8` on a card without NVFP4.
## 3. docker-compose.yml
Write this in `~/decosa/promo/`:
```yaml
services:
llm:
image: vllm/vllm-openai:v0.29.0
command: ["--model", "nvidia/Qwen3.8-27B-NVFP4", "--served-model-name", "qwen3.8-27b", "--max-model-len", "32768",
"--enable-prefix-caching"]
ports: ["127.0.0.1:8114:8000"]
volumes: ["~/.cache/huggingface:/root/.cache/huggingface"]
deploy: { resources: { reservations: { devices: [{ driver: nvidia, count: 1, capabilities: [gpu] }] } } }
healthcheck: { test: ["CMD", "curl", "-fs", "http://localhost:8000/v1/models"], interval: 30s, retries: 20 }
api:
image: ${DECOSA_REGISTRY}/decosa-api:<tag>
ports: ["127.0.0.1:8445:8445"]
environment:
DECOSA_HOST: 0.0.0.0
DECOSA_PORT: "8445"
DECOSA_DATA_DIR: /data
DECOSA_LLM_ROUTE: direct
DECOSA_LLM_URL: http://llm:8000/v1
DECOSA_LLM_MODEL: qwen3.8-27b
DECOSA_PROMO_MAX_CONCURRENT: "3"
DECOSA_PROMO_WORKERS: "6"
DECOSA_BUDGET_LLM_TOKENS: "60000"
volumes: ["decosa-data:/data"]
depends_on: { llm: { condition: service_healthy } }
healthcheck: { test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://localhost:8445/promo/info', timeout=4)"], interval: 30s, retries: 10 }
volumes:
decosa-data:
```
The api keeps its state (keys, receipts, this box's signing key) in the named volume `decosa-data`, not in a
host folder: the image runs as an unprivileged user (uid 10001), and a host folder that Docker creates is owned by
root, which stops the api with `PermissionError: [Errno 13] Permission denied: '/data/keys.sqlite'`. Then start everything: `docker compose up -d`.
Prefix caching matters: every sentence of one check sends the same reference pack first, so the server reuses it.
Reference packs above 24,000 characters are narrowed per sentence with BM25 instead of being sent whole.
On the first start the api service creates this box's Ed25519 key in the `decosa-data` volume (`/data/attest/` in the api container, mode 0600). Back it up with
`docker compose cp api:/data/attest ./attest-backup` and keep that copy private.
Never print it. Every model call on the direct route gets a receipt signed with that key (status `attested`): an
attestation by me, the operator, not a proof of computation. Never set `DECOSA_LLM_ROUTE=gateway` on this box: the
gateway route sends the copy to the Decosa API.
## 4. Smoke test
1. `curl -s localhost:8445/promo/info | jq '.findings | keys'` lists the finding codes.
2. Token: `T=$(curl -s -XPOST localhost:8445/demo/session -H 'content-type: application/json' -d '{"vertical":"promo-claims-check"}' | jq -r .token)`.
3. `curl -s localhost:8445/promo/samples | jq '.[0] | {piece, product, references: [.references[] | {title, kind, text}], claims_matrix}' > planted.json`,
then `curl -s -XPOST localhost:8445/promo/check -H "authorization: Bearer $T" -H 'content-type: application/json' -d @planted.json > res.json`.
Expect `status: "issues"`; `off_label` on the weight-loss and heart-protection sentences; `contradicted` on the
"up to 2.5 percentage points" sentence; `comparative_unbacked` on "better than any other diabetes pill"; and
`fair_balance_missing` plus `boxed_warning_missing` in `piece.findings`.
4. Do the same with `.[1]` (metformin-compliant): expect no issue-level finding (status `clean` or `checks`).
5. `jq '{packet, packet_md}' res.json | curl -s -XPOST localhost:8445/promo/verify -H 'content-type: application/json' -d @-`
must show `valid_signature`, `signed_by_this_server` and `packet_md_matches` true. Change one finding in the packet
and verify again: it must fail.
6. Stream one with `-H 'accept: text/event-stream' -N`: `ready`, then `receipt` and `claim` events, then `piece`,
`packet`, `budget` and `done`.
7. Time it. On our RTX PRO 6000, shared with other work, a piece of 7-14 sentences took 4-10 s on the direct
route on 25 Sep 2026 (up to about a minute when the card is busy). Tell me what you measure.
## 5. Point the app at the local API
Set `NEXT_PUBLIC_DECOSA_API=http://127.0.0.1:8445` in the site's `.env.local`, or call `POST /promo/check` from the
tool your team uses to route pieces to review, and attach `packet_md` and the signed packet to the review job. Contract:
`API_CONTRACT.md`, section "Promotional-claims pre-check".
Off by default. Joining serves other people's requests on this GPU; never do it on a box that holds pre-launch pieces or
unpublished study data. If I ask for it, follow the provider guide at `/provide` on the site, and do not enable it
without my explicit yes.Rules and regulations it checks againstDated, linked to the primary source; not legal advice
Regulation watch
Loading the watch status…
5 laws, rules and guidance pages cited; 5 watched nightly at the primary source. A change marks this page for a human re-check; nothing is edited automatically. What we cite and how it is watched
Technical detailsModels, where it runs, labels
In short
Last reviewed
- What it is
- Every claim in a drug, device or supplement piece checked against its references before MLR review, with the span behind it and what is wrong.
- Who it's for
- Teams in healthcare and compliance and trust.
- Where it runs
- Self-host for unpublished data; hosted for public labels
- Key numbers
- 20 / 20 Planted problems caught (all categories) (test split, n = 20)
- 2 / 2 Piece-level problems caught (fair balance, DSHEA) (test split, n = 2)
- 1 / 63 (1.6%) Unplanted sentences with an issue (false positives) (test split, n = 63)
- 3.9 s Median end-to-end run, hosted (QA sweep 2026-09-25)
- Models
- Qwen3.8-27B
- Where
- Self-host for unpublished data; hosted for public labels
- Checks
- Receipt per model call; signed packet
- Industry
- Healthcare · Compliance and trust
- Input
- Text and documents
- Output
- Notes, reports and drafts · Signed record or verdict
- Data
- Confidential business data
- Hardware
- 1× 96 GB GPU
- Licence
- Permissive (Apache-2.0, MIT)
- Runs in
- Decosa hosted · Self-host
- Built from
- Grounding · Signed record
Questions people ask
What does the Promotional-claims pre-check check?
The Promotional-claims pre-check takes a drug, device or supplement piece and its reference pack: label, cited studies, spec sheets and an optional claims matrix. Each sentence gets a support verdict with the reference span behind it, then a typed check for off-label use, comparative or superlative claims, and supplement disease or structure/function claims. The whole piece is checked for fair balance, the boxed warning and the DSHEA disclaimer. It approves nothing; the MLR committee stays responsible.
How accurate is the Promotional-claims pre-check?
On 8 held-out synthetic pieces, the Promotional-claims pre-check caught 20 of 20 planted problems and 2 of 2 piece-level problems (fair balance, DSHEA), and raised an issue on 1 of 63 unplanted sentences (1.6%). The planted problems were blatant and written by the person who built the checker, the copy is synthetic, and subtle violations are not measured yet. A person reviews every finding.
Can the pre-check catch disease claims in supplement copy?
Yes, that is one of its typed checks. A line such as "Daily magnesium lowers high blood pressure naturally" is marked an issue even when the reference partly supports it, because under DSHEA (21 CFR 101.93) treating or mitigating a disease is a claim a supplement cannot make, and structure/function claims need the 101.93(c) disclaimer. Findings come from a language model and can be wrong in both directions.
Does the pre-check assess visual fair balance and layout?
No. The Promotional-claims pre-check reads text only, so type size, placement and contrast of risk information are not assessed, and tables flattened to text can be misread. Fair balance is measured as text share plus whether the boxed warning and contraindications appear. A layout-aware check is on the wanted list and not built; a claim the references support can still be misleading in context.
Can the pre-check run on our own hardware for unpublished data?
Yes. Self-host the Promotional-claims pre-check for unpublished study data and pre-launch pieces; on the direct route nothing leaves the box. The hosted demo stores nothing: the piece and references stay in memory for the request, and the signed packet holds hashes, offsets, finding codes and receipt ids, never the copy. Use hosted with public labels and synthetic or published copy. The metformin sample took 14 model calls, about $0.02.
Ask a question or leave feedbackWe read every message and publish useful answers
Ask about Promotional-claims pre-check
We read every message. Questions, comments and our answers show here once we have reviewed and approved them.
Loading questions…