Skip to content
decosa

183 · Science and research · Healthcare · preview

What studies found

Open the toolJSON

Eval results

Scored on a held-out or test splitRun 30 Sep 2026Eval write-up (decosa-api, access required)

  • Verdict matches the systematic review's (benefit, harm or no claim)231 / 289 (80%)held outn = 28995% CI 75 to 84%. The engine it replaced, on the same pairs: 189 / 289.
  • Calls a benefit where the review found no clear difference1 / 61 (2%)held outn = 61The engine it replaced, on the same pairs: 8 / 61.
  • Harms surfaced: the verdict is Worsens, or the table carries its harm flag36 / 40 (90%)held outn = 40By the verdict alone: 28 / 40. The engine it replaced, on the same pairs: 18 / 40.
  • Harm verdict or harm flag where the review reports no harm (the price of the flag)46 / 249 (18%)held outn = 249A harm verdict alone: 6 / 249. The flag is a prompt to read a study, so it errs toward flagging.
  • Finds the benefit where the review found one56 / 82 (68%)held outn = 82The engine it replaced, on the same pairs: 56 / 82. The misses are mostly tables that come out mixed or unclear.
  • Verdict matches exactly, five ways (improves, worsens, no clear difference, mixed, unclear)180 / 282 (64%)held outn = 282Pairs where both labellers gave the same five-way label. The engine it replaced, on the same pairs: 135 / 282.
  • The review behind the label was among the studies read256 / 290 (88%)held outn = 290Cochrane reviews: 143 / 154. The engine it replaced, on the same pairs: 161 / 290.
  • One abstract read alone: verdict matches the label (benefit, harm or no claim)290 / 331 (88%)held outn = 331Harms surfaced per study: 60 / 62.

Dataset

290 supplement and outcome pairs, each with a published systematic review, none used while building the engine (83 where the review found a benefit, 40 a harm, 61 no clear difference, 79 too uncertain to tell, 27 mixed). Labels: two blind AI labellers (AI sub-agents) read each review's abstract; a pair counts only when both call it a relevant supplement review and agree on benefit / harm / no claim. Each table was limited to studies published up to the review's year. Scored once, by a rule written down before any output existed; 1 pair could not be run.

Caveats

  • The one-line verdict is not a substitute for a systematic review: read the rows and the linked papers.
  • Labels are by two blind AI sub-agents reading each review's abstract, not by clinicians or systematic reviewers; pairs where they disagreed on benefit, harm or no claim were left out.
  • Each table was limited to studies published up to the review's year, so the review itself could be found; a search today can read newer studies and say something else.
  • The harm pairs came from harm topics written down before any search; they are not a random sample of supplement questions.
  • Abstracts only; no risk-of-bias assessment.
  • Model calls used the direct route to the same weights as the gateway; cost is at list price from token counts.

Nightly smoke check

Loading the nightly status…

Result
pass
Run
30 Sep 2026
Latency, this run
n/a
p50 over passed runs
26 s
Receipts
32
Model calls
n/a
Tokens
n/a
Cost per run
$0.023

Self-host verification

Not yet verified on a fresh self-host setup.

Rehearsal bundle: what-studies-found.zip (1 KB, 14 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.

Known limits

  • Hosted: measured on production on 30 Sep 2026 with the first sample (probiotics and antibiotic-associated diarrhea) through the API, one run at a time. Well-studied pairs like the samples read more reviews and cost more than the average table.
  • Reads abstracts only; results reported only in the full paper are missed.
  • The one-line verdict is not a systematic review. On held-out pairs it matches the published review's more often than not but not always (the figures are in the eval); the misses are mostly tables that come out mixed or unclear where the review found a benefit.
  • The rubric puts the newest Cochrane review first even when it studied a narrower group than the question (the melatonin sample: a review of shift workers decides, the other reviews disagree, and the verdict is "Mixed results"). Choosing the review by who it studied is planned, not built.
  • The same search can give a different verdict on another run: the model's calls on what is relevant vary a little, and PubMed changes.
  • The harm flag errs toward flagging: on held-out pairs it also appeared on tables whose review reports no harm (the figure is in the eval). It is a prompt to read the study.
  • The band has no risk-of-bias assessment; it is our rubric over the abstracts found, not a GRADE rating.
  • A meta-analysis and trials it already pooled can sit in the same table.

Models and licences

Standard tier. Licence posture: permissive (Apache, MIT or BSD).

  • Evidence engine: PubMed search and fetch, quote and n checks in code, the wording guard, the rubric grade and the signed record (CPU)decosa-evidence engine (decosa_api/studies) with the tool's routes (decosa_api/verticals/studies)Apache-2.0 (engine); AGPL-3.0-or-later (routes)
  • Model: reads each abstract (design, n, population, direction, finding, quote), reads it a second time looking only for harm, then judges each finding against the same abstractQwen3.8-27B (NVFP4)Apache-2.0
  • Ranks the PubMed results by relevance to the supplement and outcome before any abstract is readQwen3-Reranker-4B (evidence retrieval block)Apache-2.0

All quality evidence

Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.

Lite · one 32 GB card, no reranker (1)
  • Table quality without the reranker: not measured yetno run without the reranker has been scored
Standard · one 96 GB card (measured; hosted demo) (5)
  • Held-out: the table's verdict matches the systematic review's (benefit, harm or no claim): 231 / 289 (80%)decosa-api docs/evals/what-studies-found.md, 2026-09-30; held-out split, 289 pairs, scored once
  • Held-out: calls a benefit where the review found no clear difference: 1 / 61 (2%)decosa-api docs/evals/what-studies-found.md, 2026-09-30; held-out split, 289 pairs, scored once
  • Held-out: harms surfaced by the verdict or the harm flag: 36 / 40 (90%)decosa-api docs/evals/what-studies-found.md, 2026-09-30; held-out split, 289 pairs, scored once
  • Held-out: harm verdict or flag where the review reports no harm: 46 / 249 (18%)decosa-api docs/evals/what-studies-found.md, 2026-09-30; held-out split, 289 pairs, scored once
  • Held-out: finds the benefit where the review found one: 56 / 82 (68%)decosa-api docs/evals/what-studies-found.md, 2026-09-30; held-out split, 289 pairs, scored once

How we measure · All tools