Skip to content
decosa
LabsHostedSelf-host

See what studies found for a supplement

A table of what the trials found, each row linked to its PubMed record and checked against its abstract, with the grade and its reasons.

Held-out test231 / 289 (80%)Verdicts that match the published systematic review, on held-out pairs (held-out test)
On production26 smedian on production (2026-09-30); slower when the service is busy
List price~$1.46 per 100 evidence tablesmeasured, at list price

Built on: Evidence tables from PubMed, Evidence retrieval, Grounding, Signed record

Loading the tool…

Use it your way

Use it from your codeThe hosted API with your key, and prompts to paste into a coding agent
Hosted · by Decosa

Get an API key

  • Call the what studies found API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Nothing to install; we run the models.
Self-host · your GPUs

Run it yourself, on request

  • The same open models and app, on 1x RTX PRO 6000 (96 GB) or 1x RTX 5090 (32 GB) for Qwen3.8-27B; about 8 GB for the reranker; the search, checks and grade run on CPU.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.

Build with it

Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.

Base URL
https://api.decosa.ai
Auth
Authorization: Bearer $DECOSA_API_KEY (or a demo session token)
Tool id
what-studies-found

Use the hosted API

# Decosa See what studies found for a supplement: use the hosted API

You are wiring Decosa's evidence table into this project. Given a supplement and an outcome, it searches PubMed for
trials, meta-analyses and systematic reviews and returns one row per study: PMID and PubMed link, design, n, population,
direction, the finding, a quote of at most 15 words that code found word for word in the abstract, and whether the
finding held up against its own abstract. A published rubric in code grades the body of evidence, and the studies it
left out come back with why. Each model call has its own signed receipt and the table ends in a signed record. Use only
what is listed below. If you need something else, stop and ask me.

- Base URL: `https://api.decosa.ai`
- Health check: `GET https://api.decosa.ai/healthz`.
- It reports what studies found. It is not medical advice and gives no doses: a dosing or "should I take" question is
  refused with 400. Never show a row as a recommendation, and link every row to its PubMed record.

## Auth: API key (or a demo session)
1. Preferred: an API key (`dk_…`) from "Get an API key" on the tool's page. Keep it in an environment variable,
   `DECOSA_API_KEY`, never in code. Send `Authorization: Bearer $DECOSA_API_KEY`.
2. Without a key: `POST https://api.decosa.ai/demo/session` with `{"vertical": "what-studies-found"}` returns
   `{"token", "expires_at", "budget"}`. Over a limit: HTTP 429 with `Retry-After`.
3. One table at a time per demo token (409 otherwise).

## Endpoints
- `GET /evidence/table/info` (no token): the rubric's full text and version, the wording guard, limits and data flow.
- `GET /evidence/table/samples` (no token): `{samples: [{id, title, supplement, outcome, summary}]}`.
- `POST /evidence/table` (token or key). Body: `{"supplement": "probiotics", "outcome": "antibiotic-associated diarrhea",
  "max_studies"?: 8, "designs"?: "trials" | "primary", "published_before"?: 2020, "stream"?: true}` or `{"sample_id": "probiotics-antibiotic-diarrhea"}`.
  - Limits: supplement up to 80 characters, outcome up to 120, 1 to 12 studies (8 by default).
  - `"designs": "primary"` reads primary trials only; the default also reads meta-analyses and systematic reviews.
  - With `"stream": true` it streams SSE: `search` (the PubMed term and hit counts), `candidates` (the studies picked),
    `receipt` after each model call (with its `pmid`), `row` per study read, `grade`, `result`, `budget`, `done`.
  - Without streaming: one JSON object `{supplement, outcome, headline, grade: {direction, label, claim, band, consistency,
    counts, reasons, rubric, harm_signal, safety}, rows: [...], excluded: [{pmid, title, why, harm}], search, steps,
    record, record_check, usage, ...}`. `grade.label` is the verdict to show: `Improves`, `Worsens`, `No clear difference`,
    `Unclear: evidence too weak to tell`, `Mixed results` or `Too few studies to tell`. "No clear difference" and
    "Unclear" are different verdicts: keep them apart. `grade.harm_signal` is `{level: reported | possible | none, label,
    studies: [{pmid, year, design, level, quote, counted}]}`: when the level is not `none`, show its `label`
    ("Possible harm: check the studies") with the studies, whatever the verdict says.
  - Each row: `{pmid, url, title, journal, year, design, n, n_in_abstract, k_studies, population, comparator, direction,
    finding, quote, quote_check, grounding: {verdict, ...}, harm: {level, quote, flag}, net_label, abstract_sha256,
    receipt_ids}`. `direction` is `improved`, `no_difference`, `worsened` or `mixed`; `net_label` is the row's verdict
    in words; `harm.flag` `possible` means the abstract reports a worse result that did not become the row's verdict. `grounding.verdict` is `supported` or `partial`; unsupported and
    contradicted rows are dropped and listed in `excluded`.
  - Errors: 400 missing or long fields, or a dosing or advice question; 502 PubMed did not answer; 402, 409, 429.
- `POST /record/verify` (no token) `{"record": {...}}` → `{ok, summary, bad}`.

## Example: build a table and print it (Python, `pip install httpx`)
```python
import httpx, os
API = "https://api.decosa.ai"
H = {"Authorization": f"Bearer {os.environ['DECOSA_API_KEY']}"}
r = httpx.post(f"{API}/evidence/table", headers=H, timeout=180,
               json={"supplement": "magnesium", "outcome": "sleep quality", "stream": False})
r.raise_for_status()
t = r.json()
for row in t["rows"]:
    print(row["url"], row["design"], row["n"], row["direction"], row["finding"], row["grounding"]["verdict"])
print(t["grade"]["label"], "| band:", t["grade"]["band"], t["grade"]["reasons"])
if t["grade"]["harm_signal"]["level"] != "none":
    print(t["grade"]["harm_signal"]["label"], [s["pmid"] for s in t["grade"]["harm_signal"]["studies"]])
for x in t["excluded"]:
    print("left out:", x["pmid"], x["why"])
```

## Honest limits
- It reads abstracts only, so results reported only in the full paper are missed.
- The one-line verdict is not a systematic review. On held-out pairs it matched the published review's in most tables,
  not all: show the rows, and treat the band as "our rubric over the abstracts found", not as a GRADE rating. The
  numbers are on the tool's page and in its eval.
- The rubric puts the newest Cochrane review first even when it studied a narrower group than the question. When reviews
  of different groups disagree the verdict is "Mixed results": show each row's `population`.
- The harm flag errs toward flagging. It is a prompt to read the study, not a finding of harm.
- The same search can give a different verdict on another run.
- A meta-analysis and trials it already pooled can both appear as rows.

Run it yourself (containers)

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

# Decosa See what studies found for a supplement: run it yourself (containers)

You are setting up Decosa's evidence table on this machine, so the tables, the model and the signing key are your own.
Given a supplement and an outcome it searches PubMed, reads each abstract with an open model, checks each quote and
finding against the abstract, grades the body of evidence with a published rubric in code and signs a record with this
box's own key. The only outside call is to NCBI's public PubMed service (the search words and PMIDs). Nothing is sent to
Decosa's hosted API.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/what-studies-found.zip (1 KB, 14 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py what-studies-found` (the api image carries the same bundle under /app/rehearsal/what-studies-found/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py what-studies-found --bundle what-studies-found.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "probiotics and antibiotic-associated diarrhea: the studies found a better result with probiotics", "the verdict is a benefit claim, labelled Improves", "a systematic review decides the verdict"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. GPU and Docker: `nvidia-smi` must show a card with at least 32 GB (Qwen3.8-27B NVFP4 is about 20 GB of weights plus
   KV cache). If `docker compose version` fails or the NVIDIA container toolkit is missing, install them from the official
   Docker and NVIDIA instructions after asking me.
2. Fetch the compose file:
   `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
   Read it. Keep the `api` and `llm` services, and the `retrieval` service if there is room for about 8 GB more (the
   reranker). Keep the named `/data` volume and bind every port to 127.0.0.1.
3. Settings on the `api` service: `DECOSA_NCBI_TOOL` (a name for your deployment) and `DECOSA_NCBI_EMAIL` (a contact
   address I give you), so NCBI can reach you; optionally `NCBI_API_KEY` for higher limits. Without the reranker set
   `DECOSA_STUDIES_RERANK=off` (the engine then keeps PubMed's order and says so).
4. Pull and start: `docker compose pull && docker compose up -d`.
5. Check: `GET /evidence/table/info` shows `rubric.version` and `model.route`; `GET /attest/signing-key` shows this box's
   public key. Show me the key: it is what others pin to verify my tables.
6. Smoke test: get a token with `POST /demo/session {"vertical":"what-studies-found"}` and run
   `POST /evidence/table {"sample_id": "probiotics-antibiotic-diarrhea", "stream": false}`. Expect rows whose `url` starts with
   `https://pubmed.ncbi.nlm.nih.gov/`, quotes of at most 15 words with `quote_check: "found"`, a `grade` with `reasons`,
   and `record_check.ok: true`. Then `POST /record/verify` with the `record`: `ok` must be true; change one row's
   `direction` in it and verify again: it must fail. Finally `POST /evidence/table {"supplement": "probiotics",
   "outcome": "what dose should I take"}` must return 400.
7. Report back: the public key and key id, the sample's band and reasons, and how long the table took.
Run it on your own hardwareWhat it needs, and the prompt that sets it up

Run it on your own GPU

Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.

  • CPU only, 64 GB RAMDoesn't fit

    Qwen3.8-27B (NVFP4) needs a GPU.

  • GeForce RTX 4090lite tierRuns with a smaller tier

    The standard tier does not fit: Needs about 30.6 GB of GPU memory at the smallest settings; 24 GB available. The lite tier fits with changes.

  • GeForce RTX 5090lite tierRuns with a smaller tier

    The standard tier does not fit: Needs about 38.6 GB of GPU memory at the smallest settings; 32 GB available. The lite tier fits with changes.

  • 2x GeForce RTX 5090standard tierRuns

    The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

  • L40Sstandard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.

  • H100 80 GB (SXM)standard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.

  • RTX PRO 6000 Blackwell 96 GBstandard tierRuns

    The standard tier fits (68.2 of 96 GB).

  • 2x RTX PRO 6000 Blackwell 96 GBstandard tierRuns

    The standard tier fits (68.2 of 192 GB).

  • Apple M3 Ultra (Mac Studio), 96 GBlite tierRuns with a smaller tier

    The standard tier can't be checked: Qwen3-Reranker-4B (evidence retrieval block) has no mapped Apple Silicon build The lite tier fits with changes.

  • Apple M5 Max, 64 GBlite tierRuns with a smaller tier

    The standard tier can't be checked: Qwen3-Reranker-4B (evidence retrieval block) has no mapped Apple Silicon build The lite tier fits with changes.

Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

  1. 1

    Check the GPU, Docker and the NVIDIA Container Toolkit

    The driver must see the GPU, and Docker must be able to pass it into a container.

    nvidia-smi
    docker compose version
    docker run --rm --gpus all ubuntu nvidia-smi
  2. 2

    Fetch the compose file

    One file describes the API and the language model as services.

    mkdir -p ~/decosa && cd ~/decosa
    curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml
  3. 3

    Pull and start

    The first start downloads pinned model weights, tens of gigabytes.

    docker compose pull
    docker compose up -d
  4. 4

    Check health

    Wait until the API reports ok with the language model loaded. Then point your app at the local base URL.

    curl -fsS http://localhost:<PORT>/healthz
    # {"ok": true, "llm": true, ...}
    curl -fsS -X POST http://localhost:<PORT>/demo/session \
      -H 'Content-Type: application/json' -d '{"vertical":"what-studies-found"}'

Set up with a coding agent, rehearse on mock data, then go private

  1. Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
  2. Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's expected.json. Every check must print PASS.
  3. Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
Rehearsal command
docker compose exec api python scripts/rehearse.py what-studies-found

Download the mock-data bundle (1 KB, 14 checks)expected.json

Builds the evidence table for probiotics and antibiotic-associated diarrhea from PubMed (a well-studied pair: a Cochrane review and several meta-analyses of randomised trials found fewer cases with probiotics), then the table for melatonin and sleep onset latency, where the reviews disagree (the Cochrane review read is about shift workers, the other reviews are about people with sleep problems) and the verdict is "Mixed results", then asks for a dose and must be refused. The tables must link every row to PubMed by PMID, keep quotes at 15 words or fewer, grade the body with the published rubric, carry the harm flag, and end in a signed record that verifies.

What the rehearsal checks
  • probiotics and antibiotic-associated diarrhea: the studies found a better result with probiotics
  • the verdict is a benefit claim, labelled Improves
  • a systematic review decides the verdict
  • the band is low or better (it is the certainty the best review states)
  • at least four studies are counted
  • the table carries a harm flag (none, possible or reported) computed from every study read
  • the rubric in force is the one that reads harm asymmetrically
  • every row links to PubMed
  • the record is signed and verifies
  • melatonin and sleep onset: the reviews read disagree, so the verdict is Mixed results
  • a mixed verdict claims neither a benefit nor a harm
  • the rubric's reasons say the reviews disagree
  • the table shows why: one review read is about shift workers, the others are not
  • a dose request is refused

Licence: No input files: the search runs live against public PubMed records (NCBI E-utilities). Abstracts are read for the request and not kept.

Prompt for your coding agent

# Decosa See what studies found for a supplement: run it yourself (containers)

You are setting up Decosa's evidence table on this machine, so the tables, the model and the signing key are your own.
Given a supplement and an outcome it searches PubMed, reads each abstract with an open model, checks each quote and
finding against the abstract, grades the body of evidence with a published rubric in code and signs a record with this
box's own key. The only outside call is to NCBI's public PubMed service (the search words and PMIDs). Nothing is sent to
Decosa's hosted API.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/what-studies-found.zip (1 KB, 14 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py what-studies-found` (the api image carries the same bundle under /app/rehearsal/what-studies-found/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py what-studies-found --bundle what-studies-found.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "probiotics and antibiotic-associated diarrhea: the studies found a better result with probiotics", "the verdict is a benefit claim, labelled Improves", "a systematic review decides the verdict"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. GPU and Docker: `nvidia-smi` must show a card with at least 32 GB (Qwen3.8-27B NVFP4 is about 20 GB of weights plus
   KV cache). If `docker compose version` fails or the NVIDIA container toolkit is missing, install them from the official
   Docker and NVIDIA instructions after asking me.
2. Fetch the compose file:
   `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
   Read it. Keep the `api` and `llm` services, and the `retrieval` service if there is room for about 8 GB more (the
   reranker). Keep the named `/data` volume and bind every port to 127.0.0.1.
3. Settings on the `api` service: `DECOSA_NCBI_TOOL` (a name for your deployment) and `DECOSA_NCBI_EMAIL` (a contact
   address I give you), so NCBI can reach you; optionally `NCBI_API_KEY` for higher limits. Without the reranker set
   `DECOSA_STUDIES_RERANK=off` (the engine then keeps PubMed's order and says so).
4. Pull and start: `docker compose pull && docker compose up -d`.
5. Check: `GET /evidence/table/info` shows `rubric.version` and `model.route`; `GET /attest/signing-key` shows this box's
   public key. Show me the key: it is what others pin to verify my tables.
6. Smoke test: get a token with `POST /demo/session {"vertical":"what-studies-found"}` and run
   `POST /evidence/table {"sample_id": "probiotics-antibiotic-diarrhea", "stream": false}`. Expect rows whose `url` starts with
   `https://pubmed.ncbi.nlm.nih.gov/`, quotes of at most 15 words with `quote_check: "found"`, a `grade` with `reasons`,
   and `record_check.ok: true`. Then `POST /record/verify` with the `record`: `ok` must be true; change one row's
   `direction` in it and verify again: it must fail. Finally `POST /evidence/table {"supplement": "probiotics",
   "outcome": "what dose should I take"}` must return 400.
7. Report back: the public key and key id, the sample's band and reasons, and how long the table took.

Help me customise for my hardware

Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.

Hardware

GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page

Runs with a smaller tierWhat studies found on GeForce RTX 5090: use the Lite · one 32 GB card, no reranker tier

The standard tier does not fit: Needs about 38.6 GB of GPU memory at the smallest settings; 32 GB available. The lite tier fits with changes.

What this tool's stack says about this hardware:

  • 1x RTX 5090 32 GB (fits): Estimate: Qwen3.8-27B NVFP4 needs about 20 GB of weights plus KV cache for abstracts of a few thousand tokens; run without the reranker or put it on a second card. Not run here on a 5090.

Lite · one 32 GB card, no reranker: what changesuses estimates

  • Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
Memory per component
  • Evidence engine: decosa-evidence engine (decosa_api/studies) with the tool's routes (decosa_api/verticals/studies). CPU. Runs on CPU (vram_gb 0 in stack.json).
  • Model: Qwen3.8-27B (NVFP4). ~57.6 GB (at least ~28 GB), weights 21.4 GB (from stack.json). Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)

Expected speed

Not measured.

Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.

Setup prompt for this hardware

The self-host prompt for What studies found, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.

# Set up What studies found on my hardware

Fetch https://decosa.ai/prompts/what-studies-found-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied.

## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=what-studies-found)

Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4).
Quality tier: Lite · one 32 GB card, no reranker (lite). Fit check: runs with changes, about 28 GB of 32 GB used; some memory numbers are estimates, not measurements.

First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything.

Use these components (the setup below describes the standard tier; change it to match):
- Evidence engine: decosa-evidence engine (decosa_api/studies) with the tool's routes (decosa_api/verticals/studies), CPU
- Model: Qwen3.8-27B (NVFP4) (nvidia/Qwen3.8-27B-NVFP4), 57.6 GB. Change: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

GPU placement (set each service's device and its vLLM --gpu-memory-utilization to about the share shown):
- GPU 0: Qwen3.8-27B (NVFP4) ~28 GB (88%); about 4 GB left

During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed.

The stack's own component list and compose layout: https://decosa.ai/prompts/what-studies-found-assemble.md

The proof

How we tested itEval results and end-to-end checks, hosted and self-hosted, with dates

Verified end to end

Hosted: verified 30 Sep 2026 · measured 30 Sep 2026: · p50 26 s · p95 31 s (5 runs) · ~$0.023 per run · 32 receipts

Loading the nightly status…

Self-host: not yet verified

Measured cost to run: about $1.46 per 100 evidence tables (hosted, 30 Sep 2026, partly estimated). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.

Known limits (8)
  • Hosted: measured on production on 30 Sep 2026 with the first sample (probiotics and antibiotic-associated diarrhea) through the API, one run at a time. Well-studied pairs like the samples read more reviews and cost more than the average table.
  • Reads abstracts only; results reported only in the full paper are missed.
  • The one-line verdict is not a systematic review. On held-out pairs it matches the published review's more often than not but not always (the figures are in the eval); the misses are mostly tables that come out mixed or unclear where the review found a benefit.
  • The rubric puts the newest Cochrane review first even when it studied a narrower group than the question (the melatonin sample: a review of shift workers decides, the other reviews disagree, and the verdict is "Mixed results"). Choosing the review by who it studied is planned, not built.
  • The same search can give a different verdict on another run: the model's calls on what is relevant vary a little, and PubMed changes.
  • The harm flag errs toward flagging: on held-out pairs it also appeared on tables whose review reports no harm (the figure is in the eval). It is a prompt to read the study.
  • The band has no risk-of-bias assessment; it is our rubric over the abstracts found, not a GRADE rating.
  • A meta-analysis and trials it already pooled can sit in the same table.

Eval results, nightly checks and cost per runVerify a run

How it's builtThe steps, the models and what each one checks
Hosted · by Decosa

Get an API key

  • Call the what studies found API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Nothing to install; we run the models.
Self-host · your GPUs

Run it yourself, on request

  • The same open models and app, on 1x RTX PRO 6000 (96 GB) or 1x RTX 5090 (32 GB) for Qwen3.8-27B; about 8 GB for the reranker; the search, checks and grade run on CPU.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
The open stack

A table of what the trials found for a supplement and an outcome, each row linked to its PubMed record and checked against its abstract.

Type a supplement and an outcome. It searches PubMed for trials, meta-analyses and systematic reviews, reads each abstract for design, size, population, what was found and a short quote, and checks each finding against its own abstract. A published rubric in code then grades the body of evidence by design, size and consistency. For evidence-based supplement sites, brand regulatory teams, health journalists and researchers who need a first table they can check, cite and rebuild. It reports what studies found: no advice, no doses.

Deployment
Hosted or self-host
Regulatory
Written 29 Sep 2026. This tool reports what published studies found. It is not medical advice, not a recommendation and not a dose, and it says nothing about any product. For supplement sellers: under the Dietary Supplement Health and Education Act of 1994 (DSHEA), a supplement label may carry a structure/function statement with the FDA disclaimer in 21 CFR 101.93(c) ("This statement has not been evaluated by the Food and Drug Administration. This product is not intended to diagnose, treat, cure, or prevent any disease."), but not a disease claim, which 21 CFR 101.93(g) describes as a claim to diagnose, mitigate, treat, cure or prevent disease (https://www.law.cornell.edu/cfr/text/21/101.93, read 29 Sep 2026; official text at https://www.ecfr.gov/current/title-21/chapter-I/subchapter-B/part-101/subpart-F/section-101.93). The FTC's Health Products Compliance Guidance (December 2022, https://www.ftc.gov/business-guidance/resources/health-products-compliance-guidance, read 29 Sep 2026) expects competent and reliable scientific evidence for health claims, which it says will generally need to be randomized, controlled human clinical testing, and the evidence must fit the claim. A table of what trials found is input to that judgement, not substantiation by itself and not wording you can put on a label: the product, amount and population in the trials must match, and counsel decides. The tool's wording guard (studies-guard-1) rewrites or blocks treat, cure, heal, prevent and reverse wording, personal recommendations, and doses or schedules, and a dosing question is refused. The search words go to NCBI's public E-utilities, whose usage policy applies [not re-read for this note]. PubMed abstracts can be under publisher copyright [unverified: it varies by journal], so the tool keeps PMIDs, its own structured findings, quotes of at most 15 words and a SHA-256 of each abstract, not abstract text. Model licences: Apache-2.0 (Qwen3.8-27B, Qwen3-Reranker-4B). Not legal or regulatory advice.
Architecture
Text description

A supplement and an outcome go to the evidence engine, which searches PubMed through NCBI's public E-utilities (the question as asked, then reviews and Cochrane reviews) and fetches titles and abstracts. The reranker (Qwen3-Reranker-4B, Apache-2.0) picks the most relevant studies; reviews are always read. For each study, Qwen3.8-27B (Apache-2.0) reads the abstract for design, n, population, direction, finding and a short quote, then reads it a second time looking only for harm; code checks the quote and n in the abstract and the wording guard removes treatment and dosing language; the grounding block then judges the finding against the same abstract. The rubric in code grades the body of evidence: Improves, Worsens, No clear difference, Unclear: evidence too weak to tell, Mixed results, or Too few studies to tell, with a harm flag when any study read reports a worse result. Out come the table with PubMed links, the grade with its reasons, the harm flag, the studies left out and a signed record. Each model call gets a signed receipt. Self-hosted, the engine, model and reranker stay on your machine and only the search words and PMIDs go to PubMed.

Architecture

At a glance

Data retention
The hosted service keeps the PMIDs each search returned for 7 days, keyed by a SHA-256 of the search term, so repeat searches are fast. Titles and abstracts live in memory for the request; the signed record holds each abstract's SHA-256, not its text.
What leaves the box
The search words and PMIDs go to NCBI's public PubMed service. Abstracts and search words go to the model and the reranker, which run on Decosa's hosted service (hosted) or your machine (self-host).
What it will not do
No doses, no advice on what to take, no treat, cure or prevent wording: a dosing or advice question is refused, and the guard rewrites such wording in findings.
Input
A supplement (up to 80 characters) and an outcome (up to 120), or a sample. Optional: how many studies to read, primary trials only, or only studies published before a year.
Output
A table with one row per study (PMID and link, design, n, population, direction, finding, a quote of at most 15 words, the grounding verdict, a harm flag), the grade with its reasons, the table's harm flag with the studies behind it, the studies left out and why, and a signed record verifiable at /record/verify.
Quality tiers

Pick the tier for the quality you need

Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.

  • Lite

    one 32 GB card, no reranker

    The same model and checks on one smaller card, with PubMed's own relevance order instead of the reranker. Fewer of the picked studies may be on point.

    Models
    • decosa-evidence engine (decosa_api/studies) with the tool's routes (decosa_api/verticals/studies)
    • Qwen3.8-27B (NVFP4)
    Hardware
    1x RTX 5090 32 GB (estimate)
    Quality evidence
    • Table quality without the rerankernot measured yetno run without the reranker has been scored
    Latency
    estimate: similar to standard; not measured on this card.
    Verification
    Proof: partialSelf-host onlySelf-hosted: your own box signs each model call and the record.
  • In the hosted demo

    Standard

    one 96 GB card (measured; hosted demo)

    Qwen3.8-27B reads every abstract, reads it again for harm and checks each finding; the reranker picks the studies; the rubric grades in code.

    Models
    • decosa-evidence engine (decosa_api/studies) with the tool's routes (decosa_api/verticals/studies)
    • Qwen3.8-27B (NVFP4)
    • Qwen3-Reranker-4B (evidence retrieval block)
    Hardware
    1x RTX PRO 6000 96 GB, plus about 8 GB for the reranker
    Quality evidence
    • Held-out: the table's verdict matches the systematic review's (benefit, harm or no claim)231 / 289 (80%)decosa-api docs/evals/what-studies-found.md, 2026-09-30; held-out split, 289 pairs, scored once
    • Held-out: calls a benefit where the review found no clear difference1 / 61 (2%)decosa-api docs/evals/what-studies-found.md, 2026-09-30; held-out split, 289 pairs, scored once
    • Held-out: harms surfaced by the verdict or the harm flag36 / 40 (90%)decosa-api docs/evals/what-studies-found.md, 2026-09-30; held-out split, 289 pairs, scored once
    • Held-out: harm verdict or flag where the review reports no harm46 / 249 (18%)decosa-api docs/evals/what-studies-found.md, 2026-09-30; held-out split, 289 pairs, scored once
    • Held-out: finds the benefit where the review found one56 / 82 (68%)decosa-api docs/evals/what-studies-found.md, 2026-09-30; held-out split, 289 pairs, scored once
    Latency
    measured on our server 2026-09-30, 3 tables at once, live PubMed: p50 29.9 s and p95 55.9 s per table; $0.0142 per table on average at list price (21.3 model calls). The 7 recorded sample runs took 34.47 to 72.79 s and cost $0.018512 to $0.022449 each.
    Verification
    Proof: strongHosted: each extraction, harm reading and grounding call has a gateway-signed receipt; the rerank has a model-call receipt; the record lists them.
Components

Every model in the stack

Models in this stack. Each row has a button that shows its licence, engine, verification and evidence.
ModelDetails
Evidence engine: PubMed search and fetch, quote and n checks in code, the wording guard, the rubric grade and the signed record (CPU)decosa-evidence engine (decosa_api/studies) with the tool's routes (decosa_api/verticals/studies)
0 GBProof: partial
Model: reads each abstract (design, n, population, direction, finding, quote), reads it a second time looking only for harm, then judges each finding against the same abstractQwen3.8-27B (NVFP4)nvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab)
27.8B · 20 GBProof: strongIn the hosted demo
Ranks the PubMed results by relevance to the supplement and outcome before any abstract is readQwen3-Reranker-4B (evidence retrieval block)Qwen/Qwen3-Reranker-4B on Hugging Face (opens in a new tab)
4B · about 8.1 GB (estimate)Proof: partialIn the hosted demo

Around the models

Tools, services and hardware

Tools

  • NCBI E-utilities (PubMed) (opens in a new tab)Public service of the US National Library of Medicine; PubMed records are public, abstracts may be under publisher copyright

    esearch for the PMIDs (trials, meta-analyses and systematic reviews; retractions, errata, comments, editorials and letters excluded), efetch for titles and abstracts. Identifies itself with a tool name and, if set, a contact address and API key.

  • Grounding block (tool 17)AGPL-3.0-or-later

    Judges each row's finding against its own abstract: supported, partial, unsupported or contradicted. Unsupported and contradicted rows are dropped before grading and listed as left out.

  • Rubric studies-rubric-4 and rule studies-net-2 (decosa_api/studies/grade.py, direction.py, safety.py)Apache-2.0

    Works out each study's direction in code from separate fields (the measure's polarity, the raw change, significance, the authors' conclusion and stated certainty), then grades the body: the best available review decides (the newest Cochrane review first), its stated GRADE certainty is the band, reviews that disagree give "Mixed results", and trials alone need minimum sizes. "No clear difference" and "Unclear: evidence too weak to tell" are two different verdicts. Harm is read separately: when any study read reports a worse result for the outcome, the table carries a harm flag ("Possible harm: check the studies") whatever the verdict. GET /evidence/table/info returns the full text.

  • POST /record/verifyAGPL-3.0-or-later

    Checks the signed record and names the first entry that was changed.

Services

  • decosa-api:8000
    ${DECOSA_REGISTRY}/decosa-api:<tag>

    GET /evidence/table/info and /evidence/table/samples; POST /evidence/table (SSE or JSON). Keeps the PMIDs each search returned (keyed by the search term's SHA-256, 7 days); titles and abstracts live in memory for the request.

  • vLLM:8114
    vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1

    Qwen3.8-27B NVFP4 behind our gateway (hosted) or called directly (self-host).

  • decosa-retrieval (optional):8499
    built from decosa-api services/retrieval/Dockerfile (RETRIEVAL_RERANK=qwen3-rr-4b)

    The reranker; set DECOSA_RETRIEVAL_URL. With DECOSA_STUDIES_RERANK=off the engine uses PubMed's order.

Hardware

  • 1x RTX PRO 6000 Blackwell 96 GB Fits

    Measured on our server: the eval and the recorded demo runs ran on this card, shared with other services, with the reranker on a second card.

  • 1x RTX 5090 32 GB Fits

    Estimate: Qwen3.8-27B NVFP4 needs about 20 GB of weights plus KV cache for abstracts of a few thousand tokens; run without the reranker or put it on a second card. Not run here on a 5090.

Latency per lane

  • One table (8 studies asked for, up to 5 reviews always read), held-out split, 3 tables at once, live PubMed29.9 s

    Measuredmeasured on our server 2026-09-30: p50 29.9 s, p95 55.9 s over 289 tables (decosa-api docs/evals/what-studies-found.md)

Notes

  • Reads abstracts only: a result reported only in the full paper is missed, and so are secondary outcomes.
  • A meta-analysis and the trials it pooled can both be rows in the same table; the rubric weighs syntheses more but does not remove the overlap yet.
  • The band is GRADE-inspired, not GRADE: it has no risk-of-bias assessment, so without a review's stated certainty it stops at moderate.
  • The rubric puts the newest Cochrane review first even when it studied a narrower group than you asked about (shift workers, children, pregnancy). When reviews of different groups disagree the verdict is "Mixed results", and the result shows who each review studied.
  • The harm flag errs toward flagging: it is a prompt to read a study, not a finding of harm.
  • You choose how many studies to read (the API's limits are in GET /evidence/table/info); reviews are always read, so a table can hold a few more studies than you asked for. Ask for primary trials only with designs: "primary", or limit to studies published before a year with published_before.
Assemble it

Run this exact stack on your machine

Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.

what-studies-found/assemble-prompt.md137 lines
# Assemble Decosa's "See what studies found for a supplement" on this machine

You are setting up an evidence table builder. Given a supplement and an outcome, it searches PubMed for trials,
meta-analyses and systematic reviews, has an open model read each abstract (design, n, population, direction, finding
and a short quote), checks each quote word for word and each finding against its own abstract, grades the body of
evidence with a published rubric in code, and signs a record with this box's key. Work step by step, show me each
command before you run anything with `sudo`, and stop to ask if a check fails.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/what-studies-found.zip (1 KB, 14 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py what-studies-found` (the api image carries the same bundle under /app/rehearsal/what-studies-found/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py what-studies-found --bundle what-studies-found.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "probiotics and antibiotic-associated diarrhea: the studies found a better result with probiotics", "the verdict is a benefit claim, labelled Improves", "a systematic review decides the verdict"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## 0. Ground rules and licences
- Code and models: the evidence engine `decosa_api/studies` (Apache-2.0) inside decosa-api (routes AGPL-3.0-or-later);
  Qwen3.8-27B (Apache-2.0) reads and checks the abstracts; Qwen3-Reranker-4B (Apache-2.0) in the evidence retrieval
  service ranks the studies (optional).
- What leaves this machine: the search words and PMIDs, to NCBI's public E-utilities (PubMed). Nothing else. Bind every
  port to 127.0.0.1.
- Be polite to NCBI: set `DECOSA_NCBI_TOOL` and `DECOSA_NCBI_EMAIL` to a name and a contact address I give you; add
  `NCBI_API_KEY` only if I give you one. Do not add retries or raise concurrency beyond the defaults.
- Be honest about what it does: it reports what published abstracts say. It is not medical advice, gives no doses, and
  its band is a rubric over the abstracts found, not a GRADE rating. Don't add wording that says a supplement treats,
  cures or prevents anything; the engine's guard blocks it and so should anything you build on top.

## 1. Check the machine
1. `nvidia-smi`: one GPU with at least 32 GB (Qwen3.8-27B NVFP4 needs about 20 GB of weights plus KV cache for abstracts
   of a few thousand tokens). An RTX PRO 6000 96 GB fits the model and the reranker together; on a 32 GB card run the
   reranker on a second card or turn it off. Driver 570 or newer. Blackwell cards run NVFP4; on older cards use the FP8
   weights.
2. `docker --version` and `docker compose version`. If Docker or the NVIDIA container toolkit is missing, install them
   from the official Docker and NVIDIA repositories after asking me, then run
   `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
3. Outbound HTTPS to eutils.ncbi.nlm.nih.gov and huggingface.co. Disk: about 40 GB.

## 2. Images and weights
- `${DECOSA_REGISTRY}/decosa-api:<tag>` (**publishing soon**). If the pull fails, build from source:
  `git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host>` (access required), check out a release that contains
  `decosa_api/studies/`, and `docker build -f docker/api/Dockerfile -t decosa-api:local .`
- `vllm/vllm-openai:v0.29.0` for the model; weights `nvidia/Qwen3.8-27B-NVFP4` (or `Qwen/Qwen3.8-27B-FP8`).
- The retrieval service is built from the same repository: `services/retrieval/Dockerfile`, with
  `RETRIEVAL_RERANK=qwen3-rr-4b` (weights `Qwen/Qwen3-Reranker-4B`, revision 22e683669bc0f0bd69640a1354a6d0aebcfeede5).

## 3. docker-compose.yml
Write this in `~/decosa/studies/`:

```yaml
x-health: &health { interval: 30s, timeout: 5s, retries: 20 }
services:
  llm:
    image: vllm/vllm-openai:v0.29.0
    command: ["--model", "nvidia/Qwen3.8-27B-NVFP4", "--served-model-name", "qwen3.8-27b", "--max-model-len", "32768",
              "--enable-prefix-caching", "--host", "0.0.0.0", "--port", "8000"]
    ports: ["127.0.0.1:8114:8000"]
    volumes: ["hf-cache:/root/.cache/huggingface"]
    deploy: { resources: { reservations: { devices: [{ driver: nvidia, count: 1, capabilities: [gpu] }] } } }
    healthcheck: { <<: *health, test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8000/health', timeout=4)"], start_period: 900s }
  retrieval:
    build: { context: /path/to/decosa-api, dockerfile: services/retrieval/Dockerfile }
    deploy: { resources: { reservations: { devices: [{ driver: nvidia, count: 1, capabilities: [gpu] }] } } }
    environment: { RETRIEVAL_RERANK: qwen3-rr-4b }
    volumes: [retrieval-models:/models]
    healthcheck: { <<: *health, test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8499/health', timeout=4)"], start_period: 600s }
  api:
    image: ${DECOSA_REGISTRY}/decosa-api:<tag>
    ports: ["127.0.0.1:8445:8445"]
    environment:
      DECOSA_HOST: 0.0.0.0
      DECOSA_PORT: "8445"
      DECOSA_DATA_DIR: /data
      DECOSA_LLM_ROUTE: direct                 # local model; receipts signed by this box's key ("attested")
      DECOSA_LLM_URL: http://llm:8000/v1
      DECOSA_LLM_MODEL: qwen3.8-27b
      DECOSA_RETRIEVAL_URL: http://retrieval:8499
      DECOSA_NCBI_TOOL: "<deployment name>"
      DECOSA_NCBI_EMAIL: "<contact address>"
      DECOSA_STUDIES_MAX_CONCURRENT: "3"
    volumes: ["decosa-data:/data"]
    depends_on: { llm: { condition: service_healthy }, retrieval: { condition: service_healthy } }
    healthcheck: { <<: *health, test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8445/evidence/table/info', timeout=4)"], start_period: 20s }
volumes: { hf-cache: {}, retrieval-models: {}, decosa-data: {} }
```

No room for the reranker? Drop the `retrieval` service and its `depends_on`, and set `DECOSA_STUDIES_RERANK: "off"`: the
engine then keeps PubMed's relevance order and says so in each result (`rank`).

The api keeps its state (keys, receipts, this box's signing key, the PMID cache) in the named volume `decosa-data`, not a
host folder: the image runs as uid 10001, and a host folder Docker creates is root's, which stops the api with a
PermissionError. The PMID cache (`/data/studies/pmid_cache.sqlite`, 7 days) holds PMIDs keyed by a hash of the search
term, never titles or abstracts. Start: `docker compose up -d`.

On first start the api creates this box's Ed25519 key in the volume (`/data/attest/`, mode 0600). Back it up with
`docker compose cp api:/data/attest ./attest-backup`, keep it private and never print it.

## 4. Smoke test
1. `curl -s localhost:8445/evidence/table/info | jq '{rubric: .rubric.version, guard: .guard.version, model: .model, rerank: .rerank, ncbi: .ncbi}'`
   shows `studies-rubric-4`, `studies-guard-1`, route `direct`, and `email_set: true`.
2. Token: `T=$(curl -s -XPOST localhost:8445/demo/session -H 'content-type: application/json' -d '{"vertical":"what-studies-found"}' | jq -r .token)`.
3. The sample: `curl -s -XPOST localhost:8445/evidence/table -H "authorization: Bearer $T" -H 'content-type: application/json' -d '{"sample_id":"probiotics-antibiotic-diarrhea","stream":false}' > table.json`.
   Expect every row's `url` to start with `https://pubmed.ncbi.nlm.nih.gov/`, quotes of at most 15 words with
   `quote_check: "found"` (or an empty quote with the reason), `grade.label` (the sample usually reads `Improves`) with `grade.reasons`, a `grade.harm_signal`, and
   `record_check.ok: true`. `rank` is `reranked` with the retrieval service, else says PubMed's order.
4. The guard: `curl -s -o /dev/null -w '%{http_code}' -XPOST localhost:8445/evidence/table -H "authorization: Bearer $T" -H 'content-type: application/json' -d '{"supplement":"probiotics","outcome":"what dose should I take"}'`
   must print 400.
5. `jq '{record: .record}' table.json | curl -s -XPOST localhost:8445/record/verify -H 'content-type: application/json' -d @-`
   must show `ok: true`. Change one row's direction in the record and verify again: it must fail.
6. Run the smoke module from the repository: `DECOSA_API_KEY=$T python scripts/smoke/what-studies-found.py http://127.0.0.1:8445`
   must print `"ok": true`.
7. Time it and tell me what you measure; PubMed's response time is part of it.

## 5. Point your tools at the local API
Set `NEXT_PUBLIC_DECOSA_API=http://127.0.0.1:8445` in the site's `.env.local`, or call `POST /evidence/table` from your
own workflow (for example, to rebuild a site's tables when new trials appear). Contract: `API_CONTRACT.md`, section
"See what studies found for a supplement". The engine also works as a Python library: `decosa_api.studies.table.Table`
takes a PubMed client, a model function and optional rerank and grounding functions.

Off by default. Joining serves other people's requests on this GPU. If I ask for it, follow the provider guide at
`/provide` on the site, and only with my explicit yes.
Technical detailsModels, where it runs, labels

In short

Last reviewed

What it is
A supplement evidence table in seconds: type a supplement and an outcome, and get one row per trial or review from PubMed with its design, size, population, what it found and a short quote, each checked against its own abstract, then a grade from a rubric anyone can read.
Who it's for
Evidence-based supplement sites, supplement brands' regulatory teams, health journalists and researchers who need a first table of the trials with PMIDs.
Where it runs
Hosted or self-host
Key numbers

On 289 held-out pairs the table's verdict matched the systematic review's in 231 (80%). It surfaced 36 of 40 harms and called a benefit on 1 of 61 reviews that found no clear difference. Read the rows, not only the verdict.

  • 231 / 289 (80%) Verdict matches the systematic review's (benefit, harm or no claim) (held out, n = 289)
  • 1 / 61 (2%) Calls a benefit where the review found no clear difference (held out, n = 61)
  • 36 / 40 (90%) Harms surfaced: the verdict is Worsens, or the table carries its harm flag (held out, n = 40)
  • 26.3 s Median end-to-end run, hosted (QA sweep 2026-09-30)
All results, datasets and caveats
Models
Qwen3.8-27B (reads each abstract, reads it again for harm, then checks each finding against it) · Qwen3-Reranker-4B (picks the studies, evidence retrieval block)
Where
Hosted or self-host
Checks
Receipt per model call; quotes checked word for word in code; the rubric grades in code; signed hash-chained record of the search, each study and the grade
Output
Structured data · Signed record or verdict
Data
No sensitive data
Hardware
1× 96 GB GPU
Licence
Permissive (Apache-2.0, MIT)

Questions people ask

What is in a supplement evidence table?

One row per study from PubMed: the PMID with a link, the design (randomised trial, meta-analysis, review), how many people, who they were, whether the result was better, no different or worse, the finding in numbers where the abstract gives them, and a quote of at most 15 words that code found word for word in the abstract. Each finding is also judged against its own abstract.

How accurate are the rows?

Each quote is found word for word in the abstract by code, and each finding is judged against its own abstract; rows that fail are dropped and listed as left out. On a held-out set we also compared one abstract read alone with a labelled reading of the same abstract. The figures, with their split, are in the key numbers and the eval.

Can I trust the one-line verdict?

Less than the rows. On held-out pairs the verdict matched the published systematic review's in most tables, not all; the figures are in the key numbers and the eval. The misses are mostly tables that come out mixed or unclear where the review found a benefit, and the same search can give a different verdict on another run. Read the rows and the linked papers.

Will it tell me what dose to take?

No. It reports what published studies found and gives no doses and no advice on what to take; a dosing or advice question is refused. Ask a doctor or pharmacist.

Can a supplement brand use the table to support a label claim?

Only as research input. FDA rules allow structure/function statements with a disclaimer but not disease claims (21 CFR 101.93), and the FTC expects competent and reliable scientific evidence that fits the claim. Whether the trials match your product, amount and claim is for your counsel.

How is the evidence graded?

By a published rubric in code, studies-rubric-4. The best available systematic review decides the direction (the newest Cochrane review first) and its stated GRADE certainty is the band; without one the band stops at moderate, because abstracts can't show risk of bias. Reviews that disagree give "Mixed results", and the result shows who each review studied. With no review, trials vote by design and size. "No clear difference" means the studies could tell and found little or none; "Unclear: evidence too weak to tell" means the reviews rate the certainty very low. They are two different verdicts.

What does "Possible harm: check the studies" mean?

Every abstract gets a second, separate reading that looks only for harm: a worse result for the outcome you asked about. When any study read reports one, the table carries the flag whatever the overall verdict says, and lists the studies. A row can carry its own "possible harm, check the study" flag. It is a prompt to read those studies, not a finding that the supplement is harmful, and it errs toward flagging.

Ask a question or leave feedbackWe read every message and publish useful answers
Questions & feedback

Ask about What studies found

We read every message. Questions, comments and our answers show here once we have reviewed and approved them.

Loading questions…

This is a

Plain text. Please leave out personal, patient or client data.

Shown with your message if we publish it. Leave blank to post as “A visitor”.

Nothing appears here until we have read and approved it.