Skip to content
decosa

Developers · Building blocks

Evidence tables from PubMed

Give it a supplement and an outcome. It finds the trials, meta-analyses and systematic reviews in PubMed, has an open model read each abstract into a row (design, n, population, direction, the finding, a short quote), checks each quote word for word and each finding against its own abstract, and grades the body of evidence with a rubric written in code. You get a table where every row links to its PubMed record, and a signed record of how it was built.

Measured 2026-09-30. Full eval. First use (Labs): See what studies found for a supplement. The engine is Apache-2.0, in decosa-api (decosa_api/studies); the API routes are AGPL-3.0-or-later.

Watch a real run

Loading the tool…

How it works

  1. PubMed search through NCBI's public E-utilities: the supplement and the outcome, limited to trials, meta-analyses and systematic reviews in humans, then searches for reviews and Cochrane reviews of the pair; retractions, errata, comments, editorials and letters are left out. PMIDs are cached by a hash of the search term.
  2. The evidence retrieval block's reranker (Qwen3-Reranker-4B) orders the results by relevance to the question; without it the engine keeps PubMed's order and says so. Reviews are always read, the newest Cochrane reviews first.
  3. For each study, Qwen3.8-27B reads the abstract into a row: relevant or not, design, n, population, comparator, direction, the finding and a quote. A second, separate call reads the same abstract only for harm: a worse result for the outcome, and what it says about side effects overall. Code then checks the quotes word for word and the n in the abstract, and a guard rewrites or drops treatment, cure, prevention and dosing wording.
  4. The grounding block judges each finding against its own abstract. Unsupported or contradicted rows are dropped and listed as left out, with why.
  5. The rubric (studies-rubric-4) grades the body in code: the best available review decides the direction (the newest Cochrane review first) and its stated GRADE certainty is the band; reviews that disagree give "Mixed results"; with no review, trials vote by design and size. Very low certainty is "Unclear: evidence too weak to tell", a different verdict from "No clear difference". When any study read reports a worse result, the table carries a harm flag ("Possible harm: check the studies") whatever the verdict.
  6. The server signs a hash-chained decosa.record.v1: the search, each study's fields and abstract SHA-256, the grade and every receipt id, never abstract text.

Results

The rows first: they are what a site or a reader uses. Then how the table's one-line direction and band compare with Cochrane reviews.

Verdict matches the systematic review's (benefit, harm or no claim)231 / 289 (80%) (held out). 95% CI 75 to 84%. The engine it replaced, on the same pairs: 189 / 289.
Calls a benefit where the review found no clear difference1 / 61 (2%) (held out). The engine it replaced, on the same pairs: 8 / 61.
Harms surfaced: the verdict is Worsens, or the table carries its harm flag36 / 40 (90%) (held out). By the verdict alone: 28 / 40. The engine it replaced, on the same pairs: 18 / 40.
Harm verdict or harm flag where the review reports no harm (the price of the flag)46 / 249 (18%) (held out). A harm verdict alone: 6 / 249. The flag is a prompt to read a study, so it errs toward flagging.
Finds the benefit where the review found one56 / 82 (68%) (held out). The engine it replaced, on the same pairs: 56 / 82. The misses are mostly tables that come out mixed or unclear.
Verdict matches exactly, five ways (improves, worsens, no clear difference, mixed, unclear)180 / 282 (64%) (held out). Pairs where both labellers gave the same five-way label. The engine it replaced, on the same pairs: 135 / 282.
The review behind the label was among the studies read256 / 290 (88%) (held out). Cochrane reviews: 143 / 154. The engine it replaced, on the same pairs: 161 / 290.
One abstract read alone: verdict matches the label (benefit, harm or no claim)290 / 331 (88%) (held out). Harms surfaced per study: 60 / 62.

290 supplement and outcome pairs, each with a published systematic review, none used while building the engine (83 where the review found a benefit, 40 a harm, 61 no clear difference, 79 too uncertain to tell, 27 mixed). Labels: two blind AI labellers (AI sub-agents) read each review's abstract; a pair counts only when both call it a relevant supplement review and agree on benefit / harm / no claim. Each table was limited to studies published up to the review's year. Scored once, by a rule written down before any output existed; 1 pair could not be run.

  • The one-line verdict is not a substitute for a systematic review: read the rows and the linked papers.
  • Labels are by two blind AI sub-agents reading each review's abstract, not by clinicians or systematic reviewers; pairs where they disagreed on benefit, harm or no claim were left out.
  • Each table was limited to studies published up to the review's year, so the review itself could be found; a search today can read newer studies and say something else.
  • The harm pairs came from harm topics written down before any search; they are not a random sample of supplement questions.
  • Abstracts only; no risk-of-bias assessment.
  • Model calls used the direct route to the same weights as the gateway; cost is at list price from token counts.

Speed and cost

One table (8 studies asked for, up to 5 reviews always read), held-out split, 3 tables at once, live PubMedmeasured on our server 2026-09-30: p50 29.9 s, p95 55.9 s over 289 tables (decosa-api docs/evals/what-studies-found.md)
The recorded sample runs on this page (7)34.47 to 72.79 s each, 2026-09-30; production (api.decosa.ai) through the gateway, every call receipted; cost at list price from token counts
CostMeasured cost to run: about $1.46 per 100 evidence tables (hosted, 30 Sep 2026, partly estimated)

Limits and data

Data retentionThe hosted service keeps the PMIDs each search returned for 7 days, keyed by a SHA-256 of the search term, so repeat searches are fast. Titles and abstracts live in memory for the request; the signed record holds each abstract's SHA-256, not its text.
What leaves the boxThe search words and PMIDs go to NCBI's public PubMed service. Abstracts and search words go to the model and the reranker, which run on Decosa's hosted service (hosted) or your machine (self-host).
What it will not doNo doses, no advice on what to take, no treat, cure or prevent wording: a dosing or advice question is refused, and the guard rewrites such wording in findings.
InputA supplement (up to 80 characters) and an outcome (up to 120), or a sample. Optional: how many studies to read, primary trials only, or only studies published before a year.

API

POST /evidence/table{supplement, outcome, max_studies?, designs?: trials | primary, published_before?, stream?} or {sample_id}. SSE: search, candidates, receipt, row, grade, result, budget, done; or one JSON object. Token or dk_ key for what-studies-found.
GET /evidence/table/infoThe rubric's full text and version, the wording guard, limits and data flow. No key.
GET /evidence/table/samplesThe sample pairs. No key.
POST /record/verify{record} -> ok, summary, bad: checks the signed record. No key.
decosa-api: decosa_api.studiestable.Table(pubmed, llm, supplement=..., outcome=..., rerank=..., ground=...) streams the same events; grade.grade(rows) is the rubric; search.PubMed is a polite E-utilities client with a PMID cache. Apache-2.0, usable as a library.

Where it goes next

  • See what studies found for a supplementbuilt

    The Labs tool: type a supplement and an outcome, read the table, download it as Markdown or CSV with the signed record.

  • Evidence-based supplement sitesunblocked

    Build and rebuild a site's own evidence tables from PubMed with PMIDs, instead of copying a paid compilation; a person reviews each table before it is published.

  • Pre-checking promotional claimsupgrade

    Pairs with the promotional-claims pre-check: what the trials behind a structure/function claim actually found.

  • Scoping tables for journalists and researchersunblocked

    A first pass at what the trials say, with links, before a proper review.

Licences

Evidence engine (decosa_api/studies, package name decosa-evidence)Apache-2.0; part of decosa-api, not published as a separate package yet
The tool's routes (decosa_api/verticals/studies)AGPL-3.0-or-later
Qwen3.8-27BApache-2.0
Qwen3-Reranker-4B (evidence retrieval block)Apache-2.0
PubMed recordsPublic NLM service; abstracts can be under publisher copyright, so the engine keeps PMIDs, its own fields, quotes of at most 15 words and abstract hashes, not abstracts

What it does not do

  • Give advice or doses. It reports what published studies found; a dosing or advice question is refused.
  • Full text: it reads abstracts only, so results reported only in the paper, and most secondary outcomes, are missed. Reading PubMed Central's open-access full text is the next step.
  • De-duplication: a meta-analysis and the trials it pooled can both be rows in one table.
  • Who a review studied: the rubric puts the newest Cochrane review first even when it covers a narrower group than the question (shift workers, children). Reviews of different groups then read as "Mixed results"; choosing the review by population is the next rubric change.
  • Risk of bias: the band is GRADE-inspired, not GRADE; without a review's stated certainty it stops at moderate.
  • A reviewer workflow for sites that publish the tables: today a person checks each table by hand.
All building blocks