Skip to content
decosa
LiveHostedSelf-hostMac

Check a lay summary's numbers

A draft or checked lay summary where every number is traced to the results table, wrong figures are flagged with the right one, and the check is on record.

Measured555 / 577 (96.2%)Planted errors caught in drafts (numbers, comparisons, side effects)
On production8.2 smedian on production (2026-09-26); slower when the service is busy
List price~$0.036 per trialmeasured, at list price

Built on: Numeric grounding, Grounding, Signed record, Language pack

Loading the tool…

Use it your way

Use it from your codeThe hosted API with your key, and prompts to paste into a coding agent
Hosted · by Decosa

Get an API key

  • Call the eu trial lay summary with number grounding API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Nothing to install; we run the models.
Self-host · your GPUs

Run it yourself, on request

  • The same open models and app, on Qwen3.8-27B (1× RTX 5090 32 GB or larger); the number tracing, checks and record run on CPU.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.

Build with it

Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.

Base URL
https://api.decosa.ai
Auth
Authorization: Bearer $DECOSA_API_KEY (or a demo session token)
Tool id
eu-trial-lay-summary

Use the hosted API

# Decosa EU trial lay summary: use the hosted API

You are wiring Decosa's EU trial lay summary into this project. Given a trial with results posted on ClinicalTrials.gov
it drafts the lay summary the EU Clinical Trials Regulation requires (Regulation (EU) No 536/2014, Article 37(4); the
content is set by Annex V), or checks a draft someone wrote. Every number is traced in code to a results cell, or to
arithmetic on the cells a sentence cites; comparisons are checked against the table; a result the posted analysis says
could be chance must say so; serious side effects and deaths must be given for every group; promotional and softening
words are flagged; every narrative sentence goes through a grounding judge. It returns Annex V coverage, the reading
grade, a Markdown review file and a record signed by the server, with a signed receipt for every model call. Use only
what is listed below. If you need something else, stop and ask me.

- Base URL: `https://api.decosa.ai`
- Health check: `GET https://api.decosa.ai/healthz`.
- It never says a summary is compliant or meets Annex V. Do not add that claim in your UI. The sponsor's medical writer
  reviews every flag, fills the `[Sponsor to add: ...]` gaps (EU trial number, contact details, follow-up plans,
  participants per country) and approves.
- The hosted API is for results that are already public. Unpublished results belong on a self-hosted box.

## Auth: API key (or a demo session)
1. Preferred: an API key (`dk_…`) from "Get an API key" on the tool page. Keep it in an environment variable,
   `DECOSA_API_KEY`, never in code. Send `Authorization: Bearer $DECOSA_API_KEY`.
2. Without a key: `POST https://api.decosa.ai/demo/session` with `{"vertical": "eu-trial-lay-summary"}` returns
   `{"token", "expires_at", "budget"}`. Sessions per IP are limited; over a limit you get HTTP 429 with `Retry-After`.
3. A draft needs about 14,800 generated tokens of budget, a check with grounding about 6,400 (402 otherwise). One run at
   a time per demo token (409 while one is going).

## Endpoints
- `POST /laysummary/sources` (no token, no model): `{"nct_id": "NCT04880850"}` (or `{"sample_id": ...}` or
  `{"study": <a ClinicalTrials.gov API v2 study record>}`) → the registry facts `F1..`, results cells `C1..` (table, row,
  group, value, and for counts the group size), tables `T1..`, protocol sentences `P1..` and the posted analyses, with
  whether each difference could be chance. Only the NCT number leaves for ClinicalTrials.gov.
- `POST /laysummary/draft` (token): `{"nct_id": "...", "inputs": {"eu_ct_number": "2025-521150-42-00", "sponsor_contact":
  "...", "follow_up": "...", "participants_by_country": "..."}, "stream": true}`. English only. Takes about two minutes
  (six section drafts, one repair pass, a grounding call per sentence).
- `POST /laysummary/check`: `{"draft": "<Markdown with headings>", "nct_id": "...", "language": "en", "grounding": true}`.
  Source ids in brackets (`[C12, F3]`) are optional; an uncited number is looked up in its own section's tables. With
  `"grounding": false` it makes no model call and needs no token. `language` is en, de, fr, it, nl or es: numbers
  (decimal comma) and readability work in all six; word lists, comparisons and hedges are English only.
- Answer (JSON, or SSE with `Accept: text/event-stream` / `"stream": true`: `ready`, `section`, `receipt`, `draft`,
  `sentence`, `coverage`, `readability`, `doc_flag`, `warning`, `result`, `budget`, `done`): `{mode, draft, reader_copy,
  sentences: [{i, section, text, cites, status: ok | check | fail | needs_input, numbers: [{text, status: ok | ok_uncited
  | elsewhere | mismatch | unsupported, reason, trace: {op, cells, value, why}, nearest}], flags: [{kind, severity, why}],
  grounding}], counts, coverage (Annex V's ten elements: covered | partial | missing | needs_input), readability,
  doc_flags, first_pass, repairs, review_md, record, record_check, receipts}`.
- `POST /laysummary/translate` (token): `{"draft": "<the finished English summary>", "nct_id": "..." (or "sample_id"),
  "languages": ["de", "fr", "pl"], "stream": true}` → Member-State versions (up to six of de, fr, es, it, nl, pl, pt, cs):
  each sentence translated by Hy-MT2-7B, its numbers checked against the English sentence and traced to the results
  cells in that language: `{languages: [{lang, draft, sentences: [{en, text, cites, status, preservation, numbers}], counts,
  readability}], record}`. Machine translations: a native speaker and a lay reader read each version before submission.
- `POST /record/verify` (no token): the `record` → `{ok, checks, summary}`.
- `GET /laysummary/info`, `GET /laysummary/samples` (no token): Annex V verbatim with the date it was read, the
  deadline rule, the checks, limits, and three sample trials.

## Build it like this
Show each sentence with its status, and make every number clickable to its `trace.why` (the results cell). Show `fail`
sentences before `check` ones, and never hide `[Sponsor to add]` gaps. Store the `record` next to the approved summary.

## Errors
400 bad input (the message names the field), 401/403 token, 402 budget, 409 a run already going on this demo token,
413 body over 5 MB, 429 busy (`Retry-After`), 502 ClinicalTrials.gov could not be reached or has no such study.

Run it yourself (containers)

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

# Decosa EU trial lay summary: run it yourself (containers)

You are setting up the Decosa EU trial lay summary on this machine, so unpublished trial results never leave it. It drafts
or checks the lay summary for Annex V of Regulation (EU) No 536/2014: every number traced in code to a results cell,
comparisons and "could be chance" checked against the posted analysis, side effects checked for every group, each claim
through a grounding judge, and a signed record of every check. It never says a summary is compliant; the sponsor's
medical writer approves.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/eu-trial-lay-summary.zip (17 KB, 9 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py eu-trial-lay-summary` (the api image carries the same bundle under /app/rehearsal/eu-trial-lay-summary/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py eu-trial-lay-summary --bundle eu-trial-lay-summary.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the results are read into cells with no model call, and the icodec serious side-effect count is cell C114 (22 of 291)", "the correct glargine sentence is traced to its cell and passes", "the grounding judge supports it"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
   instructions for this distribution (docs.docker.com/engine/install). Install the NVIDIA container toolkit and check
   `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
   `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
   Read it. Keep the `llm` service (Qwen3.8-27B on vLLM, with prefix caching on) and the `api` service. For the `api`
   service set `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b` and bind
   every port to 127.0.0.1. For results that are not public, also set `DECOSA_LAYSUMMARY_FETCH=0` (then send the study
   record as `study`; nothing calls ClinicalTrials.gov). Never set the gateway route for unpublished results.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check (the first start
   downloads about 20 GB of weights).
4. Check: `curl -fsS http://127.0.0.1:<PORT>/laysummary/info` lists Annex V's ten elements and the date they were read;
   `GET /attest/signing-key` shows this box's public key. Show me the key: QC and auditors pin it to verify records.
5. Smoke test: `POST /laysummary/check` with `{"sample_id": "icodec-weekly-insulin", "grounding": false, "draft": "## What
   side effects were seen\n23 out of 291 people (8%) in the Insulin Icodec group had a serious side effect. [C114]\nNo one
   in the trial died.\n"}` (no token needed without grounding). Expect the first sentence `fail` with a number `mismatch`
   whose nearest cell is C114 (22 of 291), and a `contradicts_table` document flag. Send the `record` to
   `POST /record/verify`: `ok` must be true. Then get a token (`POST /demo/session {"vertical":"eu-trial-lay-summary"}`)
   and run `POST /laysummary/draft {"sample_id": "ruxolitinib-covid"}`: expect ten sections, every number traced, and
   "could be due to chance" in the results.
6. Member-State versions (optional): add the language pack's translation model to the compose file (vLLM serving
   `tencent/Hy-MT2-7B` at revision 9b0eb4e8f001def3e5ff6469a0ac96fdb39ec223 with `--enforce-eager`, about 18 GB) and set
   `DECOSA_LANG_MT_URL=http://mt:8000/v1` on the api. Then `POST /laysummary/translate {"sample_id": "ruxolitinib-covid",
   "draft": <the draft>, "languages": ["de"]}` must return German sentences with their numbers traced.
7. Report back: the public key, the smoke results, and how long the draft took.

Off by default. Joining as a provider serves other people's requests on this GPU; never do it on a box that holds
unpublished results. If I ask for it later, follow the Provide page instead of improvising.

No NVIDIA GPU? Parts of this tool also run on an Apple Silicon Mac (MLX, 32 GB of unified memory or more): use https://decosa.ai/prompts/eu-trial-lay-summary-mac.md instead.
Run it on your own hardwareWhat it needs, and the prompt that sets it up

Run it on your own GPU

Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.

  • CPU only, 64 GB RAMlite tierRuns with a smaller tier

    The standard tier does not fit: Qwen3.8-27B (NVFP4) needs a GPU. The lite tier fits.

  • GeForce RTX 4090lite tierRuns with a smaller tier

    The standard tier does not fit: Needs about 38 GB of GPU memory at the smallest settings; 24 GB available. The lite tier fits.

  • GeForce RTX 5090lite tierRuns with a smaller tier

    The standard tier does not fit: Needs about 46 GB of GPU memory at the smallest settings; 32 GB available. The lite tier fits.

  • 2x GeForce RTX 5090standard tierRuns

    The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

  • L40Slite tierRuns with a smaller tier

    The standard tier does not fit: Needs about 51.6 GB of GPU memory at the smallest settings; 48 GB available. The lite tier fits.

  • H100 80 GB (SXM)standard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.

  • RTX PRO 6000 Blackwell 96 GBstandard tierRuns

    The standard tier fits (75.6 of 96 GB).

  • 2x RTX PRO 6000 Blackwell 96 GBstandard tierRuns

    The standard tier fits (75.6 of 192 GB).

  • Apple M3 Ultra (Mac Studio), 96 GBstandard tierRuns

    The standard tier fits with changes: Member-State versions: Hy-MT2-7B has not been run on a Mac; the fifteen languages routed to Qwen3.8-27B use the same model as drafting. Drafting and checking in English run as before.

  • Apple M5 Max, 64 GBstandard tierRuns

    The standard tier fits with changes: Member-State versions: Hy-MT2-7B has not been run on a Mac; the fifteen languages routed to Qwen3.8-27B use the same model as drafting. Drafting and checking in English run as before.

Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

  1. 1

    Check the GPU, Docker and the NVIDIA Container Toolkit

    The driver must see the GPU, and Docker must be able to pass it into a container.

    nvidia-smi
    docker compose version
    docker run --rm --gpus all ubuntu nvidia-smi
  2. 2

    Fetch the compose file

    One file describes the API and the language model as services.

    mkdir -p ~/decosa && cd ~/decosa
    curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml
  3. 3

    Pull and start

    The first start downloads pinned model weights, tens of gigabytes.

    docker compose pull
    docker compose up -d
  4. 4

    Check health

    Wait until the API reports ok with the language model loaded. Then point your app at the local base URL.

    curl -fsS http://localhost:<PORT>/healthz
    # {"ok": true, "llm": true, ...}
    curl -fsS -X POST http://localhost:<PORT>/demo/session \
      -H 'Content-Type: application/json' -d '{"vertical":"eu-trial-lay-summary"}'

Set up with a coding agent, rehearse on mock data, then go private

  1. Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
  2. Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's expected.json. Every check must print PASS.
  3. Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
Rehearsal command
docker compose exec api python scripts/rehearse.py eu-trial-lay-summary

Download the mock-data bundle (17 KB, 9 checks)expected.json

The ONWARDS 4 trial (NCT04880850, insulin icodec weekly vs insulin glargine daily) as posted on ClinicalTrials.gov, and a three-sentence side-effects section: the glargine count is right (25 of 291), the icodec count is wrong (23; the table says 22), and 'No one in the trial died' is false (the table has 2 and 1 deaths). The check must trace the right sentence to its cell and have the grounding judge support it, flag the wrong count with cell C114 as the nearest figure, flag the death claim against the table, and sign a record that verifies and fails once changed. Sources first come back with no model call.

What the rehearsal checks
  • the results are read into cells with no model call, and the icodec serious side-effect count is cell C114 (22 of 291)
  • the correct glargine sentence is traced to its cell and passes
  • the grounding judge supports it
  • the wrong icodec count (23) is flagged as a mismatch
  • with the table's own figure as the nearest
  • 'No one in the trial died' is flagged against the table
  • the signed record verifies
  • a record whose draft hash was changed no longer verifies
  • every model call has a signed receipt

Licence: Study record: ClinicalTrials.gov NCT04880850 (US National Library of Medicine registry; results are facts reported by the sponsor, reproduced unchanged). Draft: written for this bundle (CC0). Part of decosa-api, AGPL-3.0-or-later.

Prompt for your coding agent

# Decosa EU trial lay summary: run it yourself (containers)

You are setting up the Decosa EU trial lay summary on this machine, so unpublished trial results never leave it. It drafts
or checks the lay summary for Annex V of Regulation (EU) No 536/2014: every number traced in code to a results cell,
comparisons and "could be chance" checked against the posted analysis, side effects checked for every group, each claim
through a grounding judge, and a signed record of every check. It never says a summary is compliant; the sponsor's
medical writer approves.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/eu-trial-lay-summary.zip (17 KB, 9 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py eu-trial-lay-summary` (the api image carries the same bundle under /app/rehearsal/eu-trial-lay-summary/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py eu-trial-lay-summary --bundle eu-trial-lay-summary.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the results are read into cells with no model call, and the icodec serious side-effect count is cell C114 (22 of 291)", "the correct glargine sentence is traced to its cell and passes", "the grounding judge supports it"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
   instructions for this distribution (docs.docker.com/engine/install). Install the NVIDIA container toolkit and check
   `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
   `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
   Read it. Keep the `llm` service (Qwen3.8-27B on vLLM, with prefix caching on) and the `api` service. For the `api`
   service set `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b` and bind
   every port to 127.0.0.1. For results that are not public, also set `DECOSA_LAYSUMMARY_FETCH=0` (then send the study
   record as `study`; nothing calls ClinicalTrials.gov). Never set the gateway route for unpublished results.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check (the first start
   downloads about 20 GB of weights).
4. Check: `curl -fsS http://127.0.0.1:<PORT>/laysummary/info` lists Annex V's ten elements and the date they were read;
   `GET /attest/signing-key` shows this box's public key. Show me the key: QC and auditors pin it to verify records.
5. Smoke test: `POST /laysummary/check` with `{"sample_id": "icodec-weekly-insulin", "grounding": false, "draft": "## What
   side effects were seen\n23 out of 291 people (8%) in the Insulin Icodec group had a serious side effect. [C114]\nNo one
   in the trial died.\n"}` (no token needed without grounding). Expect the first sentence `fail` with a number `mismatch`
   whose nearest cell is C114 (22 of 291), and a `contradicts_table` document flag. Send the `record` to
   `POST /record/verify`: `ok` must be true. Then get a token (`POST /demo/session {"vertical":"eu-trial-lay-summary"}`)
   and run `POST /laysummary/draft {"sample_id": "ruxolitinib-covid"}`: expect ten sections, every number traced, and
   "could be due to chance" in the results.
6. Member-State versions (optional): add the language pack's translation model to the compose file (vLLM serving
   `tencent/Hy-MT2-7B` at revision 9b0eb4e8f001def3e5ff6469a0ac96fdb39ec223 with `--enforce-eager`, about 18 GB) and set
   `DECOSA_LANG_MT_URL=http://mt:8000/v1` on the api. Then `POST /laysummary/translate {"sample_id": "ruxolitinib-covid",
   "draft": <the draft>, "languages": ["de"]}` must return German sentences with their numbers traced.
7. Report back: the public key, the smoke results, and how long the draft took.

Off by default. Joining as a provider serves other people's requests on this GPU; never do it on a box that holds
unpublished results. If I ask for it later, follow the Provide page instead of improvising.

No NVIDIA GPU? Parts of this tool also run on an Apple Silicon Mac (MLX, 32 GB of unified memory or more): use https://decosa.ai/prompts/eu-trial-lay-summary-mac.md instead.

Help me customise for my hardware

Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.

Hardware

GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page

Runs with a smaller tierEU trial lay summary with number grounding on GeForce RTX 5090: use the Lite · check a draft in code, no GPU tier

The standard tier does not fit: Needs about 46 GB of GPU memory at the smallest settings; 32 GB available. The lite tier fits.

What this tool's stack says about this hardware:

  • 1x RTX 5090 32 GB (fits): Estimate: Qwen3.8-27B NVFP4 needs about 20 GB of weights plus KV cache; not run for this use case.

Lite · check a draft in code, no GPU: what changes

Nothing: it runs as listed in the stack.

Memory per component
  • Results normaliser, number tracing, compariso...: decosa-api lay summary (decosa_api/verticals/laysummary), importing numeric grounding (53) and table cells (57), the grounding judge (17) and the signed record (07). CPU. Runs on CPU (vram_gb 0 in stack.json).

Expected speed

Not measured.

Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.

Setup prompt for this hardware

The self-host prompt for EU trial lay summary with number grounding, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.

# Set up EU trial lay summary with number grounding on my hardware

Fetch https://decosa.ai/prompts/eu-trial-lay-summary-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied.

## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=eu-trial-lay-summary)

Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4).
Quality tier: Lite · check a draft in code, no GPU (lite). Fit check: runs, about 0 GB of 32 GB used.

First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything.

Use these components (the setup below describes the standard tier; change it to match):
- Results normaliser, number tracing, compariso...: decosa-api lay summary (decosa_api/verticals/laysummary), importing numeric grounding (53) and table cells (57), the grounding judge (17) and the signed record (07), CPU

During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed.

The stack's own component list and compose layout: https://decosa.ai/prompts/eu-trial-lay-summary-assemble.md

Partly on a Mac

Some parts run natively on Apple Silicon (32 GB or more); the rest needs a CUDA GPU or a hosted API. Measured speeds and what runs where

  • Member-State versions: Hy-MT2-7B has not been run on a Mac; the fifteen languages routed to Qwen3.8-27B use the same model as drafting. Drafting and checking in English run as before.

From a checkout of decosa-api, one command sets up the models and the API: scripts/mac/setup.sh

Mac prompt for your coding agent

# Decosa EU trial lay summary with number grounding: run it on this Mac (Apple Silicon, no NVIDIA GPU)

You are setting up the Decosa EU trial lay summary with number grounding on this Mac, natively on Apple Silicon. The models run on the Mac's GPU
through MLX and decosa-api runs from a git checkout with `uv`. Docker is not used for the models, because Docker on
macOS cannot reach the GPU. Nothing is sent to Decosa's hosted API.

Only part of this tool runs on a Mac (see the gaps below). The parts that do need 32 GB of unified memory or more.

Ask me before any command that needs sudo or installs software with Homebrew, and show me the command first. Never stop
or kill a process this setup did not start; if a port is taken, pick another one.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/eu-trial-lay-summary.zip (17 KB, 9 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `.venv/bin/python scripts/rehearse.py eu-trial-lay-summary` in the decosa-api checkout (the key comes from ~/.decosa-mac/api.key).
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the results are read into cells with no model call, and the icodec serious side-effect count is cell C114 (22 of 291)", "the correct glargine sentence is traced to its cell and passes", "the grounding judge supports it"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## What runs where

| Part | On an NVIDIA GPU | On this Mac | Status |
|---|---|---|---|
| Results normaliser, number tracing, comparison and hedge checks, side-effect completeness, word lists, Annex V coverage, readability, review file and signed record (no model; CPU) | Python on CPU | The same Python module, run with uv | Runs, measured |
| Drafts the six narrative sections with citations, rewrites a failed section once, and judges each sentence against the sources | NVFP4 on vLLM 0.29 (Blackwell) | MLX 4-bit (EigenLabs/Qwen3.8-27B-4bit) on mlx_lm.server 0.31.3; oMLX 0.6.1 with MTP as an option | Runs, measured |
| Member-State versions: each sentence translated, its numbers checked against the English sentence and traced to the results cells in that language: German, French, Spanish, Italian, Dutch, Polish, Portuguese, Czech | BF16 on vLLM 0.29 (Transformers backend, eager) | Not run on a Mac; the vendor publishes a GGUF build (tencent/Hy-MT2-7B-GGUF) for llama.cpp, untested here | Untested on a Mac |
| Member-State versions: the fifteen other EU languages (Swedish, Danish, Finnish, Greek, Romanian, Hungarian, Bulgarian, Croatian, Slovak, Slovenian, Lithuanian, Latvian, Estonian; Irish and Maltese as drafts) | NVFP4 on vLLM 0.29 (Blackwell) | MLX 4-bit (EigenLabs/Qwen3.8-27B-4bit) on mlx_lm.server 0.31.3; oMLX 0.6.1 with MTP as an option | Runs, measured |

Not on a Mac:
- Member-State versions: Hy-MT2-7B has not been run on a Mac; the fifteen languages routed to Qwen3.8-27B use the same model as drafting. Drafting and checking in English run as before..

## Steps
1. Check the machine: `uname -m` must print `arm64` (an M-series chip; Intel Macs cannot run MLX), and
   `sysctl -n hw.memsize` should be at least 32 GB for this tool. Check about 30 GB of free disk with
   `df -h ~`. Show me the chip (`sysctl -n machdep.cpu.brand_string`) and the memory.
2. Tools: `uv --version`. If it is missing, ask me, then `brew install uv`.
3. Code: `git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host> ~/decosa-api` (access required) and `cd ~/decosa-api`.
   Check that `scripts/mac/setup.sh` exists; if it does not, the checkout is too old: stop and tell me.
4. Start everything with one command: `scripts/mac/setup.sh`. It creates `.venv` (decosa-api)
   and `.venv-mac` (MLX, mlx-lm, mlx-audio), downloads the weights with the Hugging Face CLI (about 16 GB for the
   language model), starts the model servers and decosa-api on 127.0.0.1, and mints a local API key
   into `~/.decosa-mac/api.key` (mode 0600). The first run takes a while because of the downloads; later runs reuse them.
   If a download fails with 401 or 403, ask me for a Hugging Face token and set `HF_TOKEN`.
   The script sets up only the language model. The parts listed under "Not on a Mac" still need
   the GPU stack or the hosted API (see https://decosa.ai/prompts/eu-trial-lay-summary-assemble.md); the smoke test in step 6 will
   report them as failures. Tell me which checks passed and which need the GPU.
5. Check health: `scripts/mac/setup.sh status` shows each server, and `curl -fsS http://127.0.0.1:8445/healthz` must
   report `"llm": true`. `curl -fsS http://127.0.0.1:8445/attest/signing-key` shows this Mac's public key:
   show it to me, because it is what others pin to check the receipts and records this Mac signs.
6. Smoke test: `.venv/bin/python scripts/mac/bench_usecases.py eu-trial-lay-summary`. It runs the tool's own sample end to end
   against the local API with the local key and prints `ok`, the wall time, the model calls and the receipts.
   `ok=True` is the pass condition. If it fails, read `~/.decosa-mac/logs/*.log` and tell me what you found.
7. Point the app at it: the API is `http://127.0.0.1:8445` with `Authorization: Bearer $(cat ~/.decosa-mac/api.key)`,
   the same routes as the hosted API. To stop everything: `scripts/mac/setup.sh stop`.
8. Report back: the chip and memory, the public key, the smoke-test result and its time, and the output of
   `scripts/mac/setup.sh status`.

## Good to know
- Receipts: every model call is signed with this Mac's own Ed25519 key and names the exact MLX weights
  (`qwen3.8-27b-mlx-4bit` with a hash of the downloaded files). There is no gateway countersignature on a
  self-hosted Mac.
- The weights are a 4-bit MLX build of the same open models, not the NVFP4 build the hosted route and the published
  evals use. Expect small differences in wording and scores.
- Faster drafting: `scripts/mac/setup.sh stop && scripts/mac/setup.sh --engine omlx` serves the
  model with oMLX and multi-token prediction (about 2x faster for a single long answer, no faster for many parallel
  calls; typed judgments then use sampling because oMLX returns no log-probabilities).
- Built from the same parts as the measured tools (Qwen3.8-27B and CPU code); not run on the Mac yet.
- Measured speeds for a Mac Studio M3 Ultra and the memory each tool needs: https://decosa.ai/mac. Full details:
  `docs/self-host-mac.md` in the checkout.

The proof

How we tested itEval results and end-to-end checks, hosted and self-hosted, with dates

Verified end to end

Hosted: verified 26 Sep 2026 · measured 26 Sep 2026: · p50 8.2 s · ~$0.001 per run · 3 receipts

Loading the nightly status…

Self-host: verified 26 Sep 2026 · fresh clone, compose up, sample against local model servers

Measured cost to run: about $0.036 per trial (hosted, 26 Sep 2026). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.

A fresh clone of a decosa-api pre-release build (not yet merged), the api image built from it with DECOSA_LAYSUMMARY_FETCH=0, run against the already-running local Qwen3.8-27B vLLM on the direct route. The rehearsal bundle passed 9 of 9 checks in 0.8 s, the smoke module passed in 0.9 s with 3 attested receipts, a full draft of the ruxolitinib sample took 18.4 s (59 attested receipts, 49 of 49 numbers traced, grade 5.7) and its record verified, and an NCT number was refused with fetching off. Torn down after. Model-server startup itself not re-verified.

Known limits (5)
  • Drafts in English; Member-State versions are machine translations with their numbers checked in code, not their wording.
  • Reads ClinicalTrials.gov records; CTIS and EudraCT results tables are not read.
  • Without citations in a draft, about half of the planted wrong numbers were missed (315 of 577 caught): cite the cells or turn the grounding judge on.
  • The grounding judge can be wrong and is not fully repeatable on the shared gateway; one correct sentence was once called contradicted.
  • Coverage says which Annex V parts are present, never that a summary meets Annex V.

Eval results, nightly checks and cost per runVerify a run

How it's builtThe steps, the models and what each one checks
Hosted · by Decosa

Get an API key

  • Call the eu trial lay summary with number grounding API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Nothing to install; we run the models.
Self-host · your GPUs

Run it yourself, on request

  • The same open models and app, on Qwen3.8-27B (1× RTX 5090 32 GB or larger); the number tracing, checks and record run on CPU.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
The open stack

Draft or check the EU lay summary of a trial's results, with every number traced to the results table.

For sponsors' clinical disclosure teams, medical writers and CROs. The EU Clinical Trials Regulation requires a summary of every trial's results for laypersons (Article 37(4); the content is set by Annex V). Give it a trial with results on ClinicalTrials.gov and it drafts that summary, or give it your own draft to check. Code turns the posted results into cells with ids (participant flow, baseline, each outcome with its statistical analysis, adverse events), registry facts and protocol sentences. Qwen3.8-27B drafts the six narrative sections, each sentence citing the ids it used; code fills the identification, sponsor, follow-up and where-to-find-more sections from registry facts and leaves marked gaps for what the registry does not hold. Then every number is traced in code to a results cell, or to arithmetic on the cells a sentence cites (22 of 291 is 8%, about 1 in 13), and a wrong one is flagged with the right figure. Comparisons are checked against the table, a result the posted analysis says could be chance must say so, serious side effects and deaths must be given for every group, and promotional or softening words are flagged. A section that fails a check is rewritten once. Each narrative sentence goes through the grounding judge (vertical 17). It reports Annex V coverage and the reading grade, and writes a review file and a signed record of every check. It never says a summary is compliant: the sponsor's medical writer approves. Member-State versions: the finished English summary is translated sentence by sentence into up to six languages by the language pack (Hy-MT2-7B); each translated number is checked against the English sentence and traced again to the results cells in that language, and the record is signed per language. Domain terms are held to a term bank: the clinical-trials pack (the Clinical Trials Regulation's terms; a vehicle cream is a placebo, not a cream for cars) plus the sponsor's own reviewed entries, with the bank version in each receipt. A native speaker and a lay reader still need to read each version.

Deployment
Hosted or self-host
Regulatory
Checked 26 Sep 2026 against the primary sources (links under Tools). Regulation (EU) No 536/2014, Article 37(4): within one year of the end of a trial in all Member States concerned, whatever the outcome, the sponsor submits a summary of the results to the EU database, accompanied by a summary understandable to laypersons whose content is set out in Annex V. Annex V lists ten elements (identification; sponsor name and contact details; where, when, objectives and reasons; the population, including numbers in the Member State concerned, the Union and third countries, age and gender breakdown and eligibility; the investigational medicinal products; adverse reactions and their frequency; overall results; comments on the outcome; whether follow-up trials are foreseen; where to find more). The 6-month deadline for paediatric trials is not in Article 37 itself: the Good Lay Summary Practice (adopted by the Clinical Trials Expert Group, 9 Jul 2021) gives 12 months, 6 for paediatric studies and up to 30 for non-therapeutic phase 1, citing the EU portal specifications (EMA/42176/2014). The expert group's recommendations (v2, 22 Feb 2018) ask for no promotional content, plain language and numeracy, explain that a non-significant difference should be explained to the reader, and call a 6th-grade Flesch-Kincaid level ideal. A cross-sectional study of the 7,547 phase II-IV trials registered in CTIS by November 2025 found that of the 234 legally required to report, 116 (49.6%) fully reported results on time (Bruckner et al., medRxiv preprint, 5 Apr 2026, not peer reviewed). Annex V asks for adverse reactions; the posted tables list adverse events whatever their cause, and the draft says so. This is a drafting and checking aid, not legal or regulatory advice, and it never states that a summary meets Annex V. Model licence: Apache-2.0 (Qwen3.8-27B). ClinicalTrials.gov records are cited by NCT number; results are facts reported by sponsors.
Architecture
Text description

A trial's results from ClinicalTrials.gov (by NCT number, or the study record pasted in) and optional sponsor inputs are normalised in code into registry facts, results cells and protocol sentences, each with an id. Qwen3.8-27B drafts the six narrative sections of Annex V, each sentence citing its ids, or the sponsor's own draft is split into sentences. Code traces every number to a results cell or to arithmetic on the cited cells, checks comparisons and 'could be chance' against the posted analysis, checks side effects for every group and flags promotional or softening words; a section with a hard problem is rewritten once with the problems listed. The grounding judge checks every narrative sentence. Outputs: the lay summary with each number linked to its cell, Annex V coverage and the reading grade, a review file and a signed record that never says compliant. On the hosted route every model call gets a gateway-signed receipt. Self-hosted, everything stays on your machine.

Architecture

At a glance

Data retention
Nothing stored: the study record, the draft and the result live in memory for the request, and logs carry counts only. A registry record fetched by NCT number is cached for a day. The signed record holds hashes, statuses and receipt ids, and the text of each sentence.
What leaves the box
Hosted: every model call goes through our gateway to the GPU serving Qwen3.8-27B, and an NCT number you type is sent to ClinicalTrials.gov. Self-hosted with DECOSA_LAYSUMMARY_FETCH=0 on the direct route: nothing leaves the box.
What it will not say
That a summary is compliant or meets Annex V. It says what it checked and marks what the registry does not hold (EU trial number, contact details, participants per country, follow-up plans) for the sponsor to add.
Input formats
A ClinicalTrials.gov NCT number or API v2 study record (CTIS and EudraCT tables have to be pasted in that shape), an optional protocol synopsis (20,000 characters), and for checking, a Markdown draft of up to 16,000 characters in English, German, French, Italian, Dutch or Spanish. Member-State versions: up to six of de, fr, es, it, nl, pl, pt, cs per request.
Typical run
A full draft: dozens of model calls, a few cents at the gateway list price, a minute or more on the shared gateway (faster self-hosted on the direct route). A short check with grounding: a few calls, a fraction of a cent. Member-State versions with the optional meaning check: two more calls per sentence per language, which adds minutes for many sentences in several languages on the busy gateway.
Quality tiers

Pick the tier for the quality you need

Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.

  • In the hosted demo

    Lite

    check a draft in code, no GPU

    POST /laysummary/check with grounding off: every number traced, comparisons, hedges, side effects, word lists, Annex V coverage and readability. No drafting and no claim-by-claim grounding.

    Models
    • decosa-api lay summary (decosa_api/verticals/laysummary), importing numeric grounding (53) and table cells (57), the grounding judge (17) and the signed record (07)
    Hardware
    Any CPU
    Quality evidence
    • Planted errors caught in drafts with citations (12 held-out trials)555 / 577 (96.2%)docs/evals/eu-trial-lay-summary.md, test split, 26 Sep 2026
    • Correct sentences flagged (false flags, same drafts)9 / 812 (1.1%)docs/evals/eu-trial-lay-summary.md, test split, adjudicated by the author
    • Planted errors caught with the citations stripped315 / 577 (54.6%)docs/evals/eu-trial-lay-summary.md, test split
    Latency
    no model call; well under a second per draft on CPU (not separately timed)
    Verification
    No proof yetNo model call, so no receipts; the checks are deterministic code.
  • In the hosted demo

    Standard

    one GPU for the model (hosted demo)

    Qwen3.8-27B drafts the six narrative sections and judges each sentence; every check and the record are code. This is what the hosted demo runs. Member-State versions: Hy-MT2-7B translates each sentence, and every number is checked again in each language.

    Models
    • decosa-api lay summary (decosa_api/verticals/laysummary), importing numeric grounding (53) and table cells (57), the grounding judge (17) and the signed record (07)
    • Qwen3.8-27B (NVFP4)
    • Hy-MT2-7B (the language-pack block)
    • Qwen3.8-27B (the language-pack block's route for these languages)
    Hardware
    1x RTX PRO 6000 96 GB (measured) or 1x RTX 5090 32 GB (estimate)
    Quality evidence
    • Wrong numbers in the model's first drafts (12 held-out trials)5 / 642 (0.8%)docs/evals/eu-trial-lay-summary.md, test split
    • Wrong numbers left after the check and one repair0 / 655docs/evals/eu-trial-lay-summary.md, test split
    • Planted errors caught by the code checks (12 held-out trials)555 / 577 (96.2%)docs/evals/eu-trial-lay-summary.md, test split
    • Planted wording errors caught by the grounding judge29 / 45 (64%)docs/evals/eu-trial-lay-summary.md, test split
    • Flesch-Kincaid grade of the drafts: median (range)6.1 (4.4-8.0)docs/evals/eu-trial-lay-summary.md, test split
    • Numbers in Member-State versions traced to the results cells (12 held-out trials, 6 languages, frozen)4,002 / 4,205 (95.2%)decosa-api docs/evals/language-pack.md, 27 Sep 2026
    • Planted number errors in the translations caught (12 trials, 6 languages)2,426 / 2,432 (99.8%)decosa-api docs/evals/language-pack.md
    Latency
    measured: a full draft takes a couple of minutes on the shared gateway, longer under load; a check with grounding takes a call per narrative sentence.
    Verification
    Proof: strongEvery model call is a separate gateway call with a gateway-signed receipt; the signed record lists them all.
Components

Every model in the stack

Models in this stack. Each row has a button that shows its licence, engine, verification and evidence.
ModelDetails
Results normaliser, number tracing, comparison and hedge checks, side-effect completeness, word lists, Annex V coverage, readability, review file and signed record (no model; CPU)decosa-api lay summary (decosa_api/verticals/laysummary), importing numeric grounding (53) and table cells (57), the grounding judge (17) and the signed record (07)
0 GBProof: partial
Drafts the six narrative sections with citations, rewrites a failed section once, and judges each sentence against the sourcesQwen3.8-27B (NVFP4)nvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab)
27.8B · 20 GBProof: strongIn the hosted demo
Member-State versions: each sentence translated, its numbers checked against the English sentence and traced to the results cells in that language: German, French, Spanish, Italian, Dutch, Polish, Portuguese, CzechHy-MT2-7B (the language-pack block)tencent/Hy-MT2-7B on Hugging Face (opens in a new tab)
7.5B · 18 GBProof: partialIn the hosted demo
Member-State versions: the fifteen other EU languages (Swedish, Danish, Finnish, Greek, Romanian, Hungarian, Bulgarian, Croatian, Slovak, Slovenian, Lithuanian, Latvian, Estonian; Irish and Maltese as drafts)Qwen3.8-27B (the language-pack block's route for these languages)Qwen/Qwen3.8-27B on Hugging Face (opens in a new tab)
27B · 0 GBProof: strong

Around the models

Tools, services and hardware

Tools

Services

  • decosa-api:8445
    ${DECOSA_REGISTRY}/decosa-api:<tag>

    GET /laysummary/info, /laysummary/samples; POST /laysummary/sources, /laysummary/draft (SSE or JSON), /laysummary/check; POST /record/verify.

  • vLLM (model):8114
    vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1

    Qwen3.8-27B NVFP4 behind our gateway (hosted) or called directly (self-host).

Hardware

  • 1x RTX PRO 6000 Blackwell 96 GB Fits

    Measured on our server: the hosted demo and the eval ran through the shared gateway.

  • 1x RTX 5090 32 GB Fits

    Estimate: Qwen3.8-27B NVFP4 needs about 20 GB of weights plus KV cache; not run for this tool.

  • CPU only Fits

    The lite tier (POST /laysummary/check with grounding off: number tracing, comparisons, hedges, side effects, words, coverage, readability) needs no GPU.

Latency per lane

  • a full draft (six sections, repair, grounding), hosted gateway route115.0 s

    Measuredmeasured on our server 2026-09-26: median over the 12 test trials, 88-289 s, 3 drafts in parallel on a gateway shared with other workloads; the 5 test2 trials took 219-363 s under heavier load

  • check a short draft with grounding, hosted gateway route8.2 s

    Measuredmeasured on our server 2026-09-26: the three-sentence smoke check took 5.0-9.1 s on the shared gateway (3 grounding calls) and 0.9 s self-hosted on the direct route

  • check a draft without the model (number tracing, comparisons, side effects, coverage, readability)n/a

    Measuredno model call; well under a second on CPU (not separately timed)

Notes

  • On 12 held-out trials (577 planted errors in drafts that had passed), the code checks caught 555 (96.2%) with citations kept, and flagged 9 of 812 correct sentences (1.1%). Every miss was a comparison or side-effect wording pattern or a loose number rule; each was fixed with a test, and 5 fresh trials then gave 212 of 217 (97.7%).
  • The model's first drafts put 5 wrong or untraceable numbers among 642 (0.8%) on the test trials; after the check and one repair pass, none was left.
  • Without citations (a sponsor's own draft), the number check falls back to the tables of the sentence's section and caught 315 of 577 (54.6%): cite the cells, or turn the grounding judge on.
  • Every test draft read at Flesch-Kincaid grade 4.4 to 8.0 (median 6.1). Annex V elements 1, 2, 4 and 9 always need the sponsor: the registry does not hold the EU trial number, contact details, participants per Member State or follow-up plans, so the draft marks those gaps instead of inventing them.
  • The grounding judge caught real slips the number trace cannot: a correct number attached to the wrong outcome, a median called an average, 'took' for 'analysed'. It also refuses some plain definitions (a placebo is a dummy treatment); a glossary source was added for that after the eval.
  • Everything was measured on ClinicalTrials.gov records with drafts from our own model. It has not been compared with published lay summaries or tested with lay readers.
Assemble it

Run this exact stack on your machine

Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.

eu-trial-lay-summary/assemble-prompt.md135 lines
# Assemble the Decosa EU trial lay summary on this machine

You are setting up a drafting and checking aid for clinical disclosure teams and medical writers. Given a trial's posted
results (a ClinicalTrials.gov record) it drafts the lay summary that Regulation (EU) No 536/2014 requires (Article 37(4);
content in Annex V), or checks a draft someone wrote: every number traced in code to a results cell, comparisons and
"could be due to chance" checked against the posted analysis, side effects checked for every group, promotional and
softening words flagged, each narrative sentence through a grounding judge, Annex V coverage and the reading grade, and
a record signed by this box's own key. It never says a summary meets Annex V; the sponsor's medical writer approves.
Work step by step, show me each command before you run anything with `sudo`, and stop to ask if a check fails.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/eu-trial-lay-summary.zip (17 KB, 9 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py eu-trial-lay-summary` (the api image carries the same bundle under /app/rehearsal/eu-trial-lay-summary/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py eu-trial-lay-summary --bundle eu-trial-lay-summary.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the results are read into cells with no model call, and the icodec serious side-effect count is cell C114 (22 of 291)", "the correct glargine sentence is traced to its cell and passes", "the grounding judge supports it"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## 0. Ground rules and licences
- Model: Qwen3.8-27B (Apache-2.0), used to draft the six narrative sections, rewrite a section once when a check fails,
  and judge each sentence against the sources. The normaliser, the number tracing, the checks, the coverage, the
  readability and the signed record are decosa-api (AGPL-3.0-or-later) and need no GPU of their own.
- Results are confidential until they are posted. For unpublished results keep everything on this machine: bind every
  port to 127.0.0.1 and set `DECOSA_LAYSUMMARY_FETCH=0`, so nothing calls ClinicalTrials.gov (send the study record as
  `study` instead of an NCT number). The service stores nothing but a fetched public record (one day, in the data
  volume); logs carry counts only. Keep it that way.
- Do not add a "compliant" or "meets Annex V" badge anywhere. The record says what was checked.

## 1. Check the machine
1. `nvidia-smi`: one GPU with at least 32 GB (Qwen3.8-27B NVFP4 needs about 20 GB of weights plus KV cache; an RTX PRO
   6000 96 GB is what we measured on; an RTX 5090 32 GB should fit but we have not run this tool on one). Driver
   570 or newer. Blackwell cards run NVFP4; on older cards use `Qwen/Qwen3.8-27B-FP8`.
2. `docker --version` and `docker compose version`. If Docker or the NVIDIA container toolkit is missing, install them
   from the official Docker and NVIDIA repositories after asking me, then run
   `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
3. Disk: about 30 GB free.

## 2. Images and weights
- `${DECOSA_REGISTRY}/decosa-api:<tag>` (**publishing soon**). If the pull fails, build from source:
  `git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host>` (access required), check out the newest release tag that
  contains `decosa_api/verticals/laysummary/` (`main` until one does), and build `docker/api/Dockerfile`.
- `vllm/vllm-openai:v0.29.0` for the model; weights `nvidia/Qwen3.8-27B-NVFP4` (revision
  `482ca0f3832238542f8f5295dde86b5f22711d80`), or `Qwen/Qwen3.8-27B-FP8` on a card without NVFP4.
- Annex V (verbatim, read 26 Sep 2026) and three sample trials ship inside the image
  (`decosa_api/verticals/laysummary/`); `GET /laysummary/info` shows the text and the sources' links.

## 3. docker-compose.yml
Write this in `~/decosa/laysummary/`:

```yaml
services:
  llm:
    image: vllm/vllm-openai:v0.29.0
    command: ["--model", "nvidia/Qwen3.8-27B-NVFP4", "--served-model-name", "qwen3.8-27b", "--max-model-len", "32768",
              "--enable-prefix-caching"]
    ports: ["127.0.0.1:8114:8000"]
    volumes: ["~/.cache/huggingface:/root/.cache/huggingface"]
    deploy: { resources: { reservations: { devices: [{ driver: nvidia, count: 1, capabilities: [gpu] }] } } }
    healthcheck: { test: ["CMD", "curl", "-fs", "http://localhost:8000/v1/models"], interval: 30s, retries: 20 }
  api:
    image: ${DECOSA_REGISTRY}/decosa-api:<tag>
    ports: ["127.0.0.1:8445:8445"]
    environment:
      DECOSA_HOST: 0.0.0.0
      DECOSA_PORT: "8445"
      DECOSA_DATA_DIR: /data
      DECOSA_LLM_ROUTE: direct
      DECOSA_LLM_URL: http://llm:8000/v1
      DECOSA_LLM_MODEL: qwen3.8-27b
      DECOSA_LAYSUMMARY_FETCH: "1"        # "0" for unpublished results: never call ClinicalTrials.gov
      DECOSA_LAYSUMMARY_MAX_CONCURRENT: "3"
      DECOSA_LAYSUMMARY_WORKERS: "4"
      DECOSA_BUDGET_LLM_TOKENS: "60000"
    volumes: ["decosa-data:/data"]
    depends_on: { llm: { condition: service_healthy } }
    healthcheck: { test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://localhost:8445/laysummary/info', timeout=4)"], interval: 30s, retries: 10 }
volumes:
  decosa-data:
```

The api keeps its state (keys, receipts, this box's signing key, the one-day registry cache) in the named volume
`decosa-data`, not in a host folder: the image runs as an unprivileged user (uid 10001), and a host folder that Docker
creates is owned by root, which stops the api with `PermissionError`. Then start everything: `docker compose up -d`.

On the first start the api creates this box's Ed25519 key in the volume (`/data/attest/`, mode 0600). Back it up with
`docker compose cp api:/data/attest ./attest-backup` and keep that copy private. Never print it. Every model call on the
direct route gets a receipt signed with that key (status `attested`): an attestation by me, the operator, not a proof of
computation. Never set `DECOSA_LLM_ROUTE=gateway` for unpublished results: that sends them to the Decosa API.

## 4. Smoke test
1. `curl -s localhost:8445/laysummary/info | jq '{n: (.annex_v | length), read: .rules_checked, fetch: .limits}'` shows
   10 elements read on 2026-09-26.
2. No model, no token: send a side-effects section with one wrong count to the check endpoint.
   ```
   curl -s -XPOST localhost:8445/laysummary/check -H 'content-type: application/json' -d '{"sample_id":"icodec-weekly-insulin","grounding":false,
     "draft":"## What side effects were seen\n23 out of 291 people (8%) in the Insulin Icodec group had a serious side effect. [C114]\nNo one in the trial died.\n"}' > chk.json
   jq '.sentences[0].numbers[0] | {status, nearest}' chk.json; jq '[.doc_flags[].kind]' chk.json
   ```
   Expect `mismatch` with nearest cell `C114` (22 of 291), and `contradicts_table` among the document flags.
3. `jq .record chk.json | jq '{record: .}' | curl -s -XPOST localhost:8445/record/verify -H 'content-type: application/json' -d @- | jq .ok`
   must print `true`. Change the text of one `sentence` entry and verify again: it must fail.
4. Token: `T=$(curl -s -XPOST localhost:8445/demo/session -H 'content-type: application/json' -d '{"vertical":"eu-trial-lay-summary"}' | jq -r .token)`.
5. `curl -s -XPOST localhost:8445/laysummary/draft -H "authorization: Bearer $T" -H 'content-type: application/json' -d '{"sample_id":"ruxolitinib-covid"}' > draft.json`,
   then `jq '{counts, grade: .readability.grade, coverage: [.coverage[] | .status]}' draft.json`. Expect ten sections,
   `numbers_wrong` 0 or close to it (read any that are not), a reading grade around 6, "could be due to chance" in the
   results section (the posted odds ratio's interval includes 1), and `needs_input` for the EU trial number, contact
   details, participants per country and follow-up plans (the registry does not hold them).
6. If decosa-api's source is at hand, `python scripts/rehearse.py eu-trial-lay-summary --base-url http://127.0.0.1:8445`
   runs the check steps and prints PASS or FAIL per property (9 checks).
7. Time it and tell me what you measure. A draft took 2 to 5 minutes through the shared hosted gateway (about 60 model
   calls); expect it to be much faster on a GPU of your own.

## 5. Point the app at the local API
Set `NEXT_PUBLIC_DECOSA_API=http://127.0.0.1:8445` in the site's `.env.local`, or call `POST /laysummary/draft` or
`/laysummary/check` from your disclosure workflow and file `review_md` and `record` with the summary. Contract:
`API_CONTRACT.md`, section "Changes (eu-trial-lay-summary, 2026-09-26)".

Off by default. Joining serves other people's requests on this GPU; never do it on a box that holds unpublished results.
If I ask for it, follow the provider guide at `/provide` on the site, and do not enable it without my explicit yes.
Rules and regulations it checks againstDated, linked to the primary source; not legal advice

Regulation watch

Loading the watch status…

3 laws, rules and guidance pages cited; 3 watched nightly at the primary source. A change marks this page for a human re-check; nothing is edited automatically. What we cite and how it is watched

Technical detailsModels, where it runs, labels

In short

Last reviewed

What it is
Draft or check the EU lay summary of a trial's results, with every number traced to the results table.
Who it's for
Clinical disclosure and transparency teams at sponsors, medical writers and CROs who write or review lay summaries for the EU database.
Where it runs
Hosted or self-host; self-host for results that are not public yet
Key numbers
  • 212 / 217 (97.7%) Planted errors caught after fixes, fresh trials (held out, n = 217)
  • 555 / 577 (96.2%) Planted errors caught, drafts with citations (code checks) (test split, n = 577)
  • 9 / 812 (1.1%) Correct sentences flagged (false flags), code checks (test split, n = 812)
  • 8.2 s Median end-to-end run, hosted (QA sweep 2026-09-26)
All results, datasets and caveats
Models
Qwen3.8-27B drafts the sections and judges each claim; number tracing, comparisons, hedges, side-effect checks, coverage and readability are plain code
Where
Hosted or self-host; self-host for results that are not public yet
Checks
Receipt per model call; every number traced to a results cell id; signed hash-chained record of the sources, the draft and every check
Output
Notes, reports and drafts · Signed record or verdict
Data
Confidential business data
Hardware
1× 96 GB GPU
Licence
Permissive (Apache-2.0, MIT)

Questions people ask

What does it check in a lay summary?

Every number against the posted results table (a cell, or a percentage or frequency computed from the cells the sentence cites), whether 'more' or 'fewer' matches the table, whether a result the analysis says could be chance is said to be, that serious side effects and deaths are given for every group, promotional or softening words, each claim through a grounding judge, Annex V's ten elements and the reading grade.

How accurate is the number check?

On 12 held-out trials it caught 555 of 577 planted errors (96.2%) in drafts with citations and flagged 9 of 812 correct sentences. Without citations it caught 315 of 577, so cite the cells or turn the grounding judge on. These are errors we planted in our own drafts, not summaries written by people.

Will it tell me my lay summary is compliant with Annex V?

No. It reports which of Annex V's ten elements it found and what it checked, and marks what the registry does not hold (the EU trial number, contact details, participants per Member State, follow-up plans). The sponsor's medical writer decides and approves.

Where do the results come from?

From ClinicalTrials.gov, by NCT number: the participant flow, baseline, each outcome with its statistical analysis, and the adverse event tables. Only the NCT number is sent. CTIS and EudraCT tables are not read yet; they have to be pasted as a study record.

Can I use it for results that are not public yet?

Yes, self-hosted: it runs on one GPU with Qwen3.8-27B, and with DECOSA_LAYSUMMARY_FETCH=0 nothing leaves the box. The hosted demo is meant for trials whose results are already public.

Which languages does it support?

It drafts in English and translates the finished summary into any EU official language (up to six at once): Hy-MT2-7B for German, French, Spanish, Italian, Dutch, Polish, Portuguese and Czech, Qwen3.8-27B for the other fifteen; Irish and Maltese come back marked "draft, needs a reviewer". Every translated number is checked against the English sentence and traced to the results cells in that language; on 12 trials 2,426 of 2,432 planted number errors were caught. Wording is not checked, so a native speaker should read each version.

Ask a question or leave feedbackWe read every message and publish useful answers
Questions & feedback

Ask about EU trial lay summary with number grounding

We read every message. Questions, comments and our answers show here once we have reviewed and approved them.

Loading questions…

This is a

Plain text. Please leave out personal, patient or client data.

Shown with your message if we publish it. Leave blank to post as “A visitor”.

Nothing appears here until we have read and approved it.