Evidence tables from PubMed
Give it a supplement and an outcome. It finds the trials, meta-analyses and systematic reviews in PubMed, has an open model read each abstract into a row (design, n, population, direction, the finding, a short quote), checks each quote word for word and each finding against its own abstract, and grades the body of evidence with a rubric written in code. You get a table where every row links to its PubMed record, and a signed record of how it was built.
Measured 2026-09-30. Full eval. First use (Labs): See what studies found for a supplement. The engine is Apache-2.0, in decosa-api (decosa_api/studies); the API routes are AGPL-3.0-or-later.
Watch a real run
Loading the tool…
How it works
- PubMed search through NCBI's public E-utilities: the supplement and the outcome, limited to trials, meta-analyses and systematic reviews in humans, then searches for reviews and Cochrane reviews of the pair; retractions, errata, comments, editorials and letters are left out. PMIDs are cached by a hash of the search term.
- The evidence retrieval block's reranker (Qwen3-Reranker-4B) orders the results by relevance to the question; without it the engine keeps PubMed's order and says so. Reviews are always read, the newest Cochrane reviews first.
- For each study, Qwen3.8-27B reads the abstract into a row: relevant or not, design, n, population, comparator, direction, the finding and a quote. A second, separate call reads the same abstract only for harm: a worse result for the outcome, and what it says about side effects overall. Code then checks the quotes word for word and the n in the abstract, and a guard rewrites or drops treatment, cure, prevention and dosing wording.
- The grounding block judges each finding against its own abstract. Unsupported or contradicted rows are dropped and listed as left out, with why.
- The rubric (studies-rubric-4) grades the body in code: the best available review decides the direction (the newest Cochrane review first) and its stated GRADE certainty is the band; reviews that disagree give "Mixed results"; with no review, trials vote by design and size. Very low certainty is "Unclear: evidence too weak to tell", a different verdict from "No clear difference". When any study read reports a worse result, the table carries a harm flag ("Possible harm: check the studies") whatever the verdict.
- The server signs a hash-chained decosa.record.v1: the search, each study's fields and abstract SHA-256, the grade and every receipt id, never abstract text.
Results
The rows first: they are what a site or a reader uses. Then how the table's one-line direction and band compare with Cochrane reviews.
| Verdict matches the systematic review's (benefit, harm or no claim) | 231 / 289 (80%) (held out). 95% CI 75 to 84%. The engine it replaced, on the same pairs: 189 / 289. |
| Calls a benefit where the review found no clear difference | 1 / 61 (2%) (held out). The engine it replaced, on the same pairs: 8 / 61. |
| Harms surfaced: the verdict is Worsens, or the table carries its harm flag | 36 / 40 (90%) (held out). By the verdict alone: 28 / 40. The engine it replaced, on the same pairs: 18 / 40. |
| Harm verdict or harm flag where the review reports no harm (the price of the flag) | 46 / 249 (18%) (held out). A harm verdict alone: 6 / 249. The flag is a prompt to read a study, so it errs toward flagging. |
| Finds the benefit where the review found one | 56 / 82 (68%) (held out). The engine it replaced, on the same pairs: 56 / 82. The misses are mostly tables that come out mixed or unclear. |
| Verdict matches exactly, five ways (improves, worsens, no clear difference, mixed, unclear) | 180 / 282 (64%) (held out). Pairs where both labellers gave the same five-way label. The engine it replaced, on the same pairs: 135 / 282. |
| The review behind the label was among the studies read | 256 / 290 (88%) (held out). Cochrane reviews: 143 / 154. The engine it replaced, on the same pairs: 161 / 290. |
| One abstract read alone: verdict matches the label (benefit, harm or no claim) | 290 / 331 (88%) (held out). Harms surfaced per study: 60 / 62. |
290 supplement and outcome pairs, each with a published systematic review, none used while building the engine (83 where the review found a benefit, 40 a harm, 61 no clear difference, 79 too uncertain to tell, 27 mixed). Labels: two blind AI labellers (AI sub-agents) read each review's abstract; a pair counts only when both call it a relevant supplement review and agree on benefit / harm / no claim. Each table was limited to studies published up to the review's year. Scored once, by a rule written down before any output existed; 1 pair could not be run.
- The one-line verdict is not a substitute for a systematic review: read the rows and the linked papers.
- Labels are by two blind AI sub-agents reading each review's abstract, not by clinicians or systematic reviewers; pairs where they disagreed on benefit, harm or no claim were left out.
- Each table was limited to studies published up to the review's year, so the review itself could be found; a search today can read newer studies and say something else.
- The harm pairs came from harm topics written down before any search; they are not a random sample of supplement questions.
- Abstracts only; no risk-of-bias assessment.
- Model calls used the direct route to the same weights as the gateway; cost is at list price from token counts.
Speed and cost
| One table (8 studies asked for, up to 5 reviews always read), held-out split, 3 tables at once, live PubMed | measured on our server 2026-09-30: p50 29.9 s, p95 55.9 s over 289 tables (decosa-api docs/evals/what-studies-found.md) |
| The recorded sample runs on this page (7) | 34.47 to 72.79 s each, 2026-09-30; production (api.decosa.ai) through the gateway, every call receipted; cost at list price from token counts |
| Cost | Measured cost to run: about $1.46 per 100 evidence tables (hosted, 30 Sep 2026, partly estimated) |
Limits and data
| Data retention | The hosted service keeps the PMIDs each search returned for 7 days, keyed by a SHA-256 of the search term, so repeat searches are fast. Titles and abstracts live in memory for the request; the signed record holds each abstract's SHA-256, not its text. |
| What leaves the box | The search words and PMIDs go to NCBI's public PubMed service. Abstracts and search words go to the model and the reranker, which run on Decosa's hosted service (hosted) or your machine (self-host). |
| What it will not do | No doses, no advice on what to take, no treat, cure or prevent wording: a dosing or advice question is refused, and the guard rewrites such wording in findings. |
| Input | A supplement (up to 80 characters) and an outcome (up to 120), or a sample. Optional: how many studies to read, primary trials only, or only studies published before a year. |
API
| POST /evidence/table | {supplement, outcome, max_studies?, designs?: trials | primary, published_before?, stream?} or {sample_id}. SSE: search, candidates, receipt, row, grade, result, budget, done; or one JSON object. Token or dk_ key for what-studies-found. |
| GET /evidence/table/info | The rubric's full text and version, the wording guard, limits and data flow. No key. |
| GET /evidence/table/samples | The sample pairs. No key. |
| POST /record/verify | {record} -> ok, summary, bad: checks the signed record. No key. |
| decosa-api: decosa_api.studies | table.Table(pubmed, llm, supplement=..., outcome=..., rerank=..., ground=...) streams the same events; grade.grade(rows) is the rubric; search.PubMed is a polite E-utilities client with a PMID cache. Apache-2.0, usable as a library. |
Where it goes next
See what studies found for a supplementbuilt
The Labs tool: type a supplement and an outcome, read the table, download it as Markdown or CSV with the signed record.
Evidence-based supplement sitesunblocked
Build and rebuild a site's own evidence tables from PubMed with PMIDs, instead of copying a paid compilation; a person reviews each table before it is published.
Pre-checking promotional claimsupgrade
Pairs with the promotional-claims pre-check: what the trials behind a structure/function claim actually found.
Scoping tables for journalists and researchersunblocked
A first pass at what the trials say, with links, before a proper review.
Licences
| Evidence engine (decosa_api/studies, package name decosa-evidence) | Apache-2.0; part of decosa-api, not published as a separate package yet |
| The tool's routes (decosa_api/verticals/studies) | AGPL-3.0-or-later |
| Qwen3.8-27B | Apache-2.0 |
| Qwen3-Reranker-4B (evidence retrieval block) | Apache-2.0 |
| PubMed records | Public NLM service; abstracts can be under publisher copyright, so the engine keeps PMIDs, its own fields, quotes of at most 15 words and abstract hashes, not abstracts |
What it does not do
- Give advice or doses. It reports what published studies found; a dosing or advice question is refused.
- Full text: it reads abstracts only, so results reported only in the paper, and most secondary outcomes, are missed. Reading PubMed Central's open-access full text is the next step.
- De-duplication: a meta-analysis and the trials it pooled can both be rows in one table.
- Who a review studied: the rubric puts the newest Cochrane review first even when it covers a narrower group than the question (shift workers, children). Reviews of different groups then read as "Mixed results"; choosing the review by population is the next rubric change.
- Risk of bias: the band is GRADE-inspired, not GRADE; without a review's stated certainty it stops at moderate.
- A reviewer workflow for sites that publish the tables: today a person checks each table by hand.