183 · Science and research · Healthcare · preview
What studies found
Eval results
Scored on a held-out or test splitRun 30 Sep 2026Eval write-up (decosa-api, access required)
- Verdict matches the systematic review's (benefit, harm or no claim)231 / 289 (80%)held outn = 28995% CI 75 to 84%. The engine it replaced, on the same pairs: 189 / 289.
- Calls a benefit where the review found no clear difference1 / 61 (2%)held outn = 61The engine it replaced, on the same pairs: 8 / 61.
- Harms surfaced: the verdict is Worsens, or the table carries its harm flag36 / 40 (90%)held outn = 40By the verdict alone: 28 / 40. The engine it replaced, on the same pairs: 18 / 40.
- Harm verdict or harm flag where the review reports no harm (the price of the flag)46 / 249 (18%)held outn = 249A harm verdict alone: 6 / 249. The flag is a prompt to read a study, so it errs toward flagging.
- Finds the benefit where the review found one56 / 82 (68%)held outn = 82The engine it replaced, on the same pairs: 56 / 82. The misses are mostly tables that come out mixed or unclear.
- Verdict matches exactly, five ways (improves, worsens, no clear difference, mixed, unclear)180 / 282 (64%)held outn = 282Pairs where both labellers gave the same five-way label. The engine it replaced, on the same pairs: 135 / 282.
- The review behind the label was among the studies read256 / 290 (88%)held outn = 290Cochrane reviews: 143 / 154. The engine it replaced, on the same pairs: 161 / 290.
- One abstract read alone: verdict matches the label (benefit, harm or no claim)290 / 331 (88%)held outn = 331Harms surfaced per study: 60 / 62.
Dataset
290 supplement and outcome pairs, each with a published systematic review, none used while building the engine (83 where the review found a benefit, 40 a harm, 61 no clear difference, 79 too uncertain to tell, 27 mixed). Labels: two blind AI labellers (AI sub-agents) read each review's abstract; a pair counts only when both call it a relevant supplement review and agree on benefit / harm / no claim. Each table was limited to studies published up to the review's year. Scored once, by a rule written down before any output existed; 1 pair could not be run.
Caveats
- The one-line verdict is not a substitute for a systematic review: read the rows and the linked papers.
- Labels are by two blind AI sub-agents reading each review's abstract, not by clinicians or systematic reviewers; pairs where they disagreed on benefit, harm or no claim were left out.
- Each table was limited to studies published up to the review's year, so the review itself could be found; a search today can read newer studies and say something else.
- The harm pairs came from harm topics written down before any search; they are not a random sample of supplement questions.
- Abstracts only; no risk-of-bias assessment.
- Model calls used the direct route to the same weights as the gateway; cost is at list price from token counts.
Nightly smoke check
Loading the nightly status…
- Result
- pass
- Run
- 30 Sep 2026
- Latency, this run
- n/a
- p50 over passed runs
- 26 s
- Receipts
- 32
- Model calls
- n/a
- Tokens
- n/a
- Cost per run
- $0.023
Self-host verification
Not yet verified on a fresh self-host setup.
Rehearsal bundle: what-studies-found.zip (1 KB, 14 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.
Known limits
- Hosted: measured on production on 30 Sep 2026 with the first sample (probiotics and antibiotic-associated diarrhea) through the API, one run at a time. Well-studied pairs like the samples read more reviews and cost more than the average table.
- Reads abstracts only; results reported only in the full paper are missed.
- The one-line verdict is not a systematic review. On held-out pairs it matches the published review's more often than not but not always (the figures are in the eval); the misses are mostly tables that come out mixed or unclear where the review found a benefit.
- The rubric puts the newest Cochrane review first even when it studied a narrower group than the question (the melatonin sample: a review of shift workers decides, the other reviews disagree, and the verdict is "Mixed results"). Choosing the review by who it studied is planned, not built.
- The same search can give a different verdict on another run: the model's calls on what is relevant vary a little, and PubMed changes.
- The harm flag errs toward flagging: on held-out pairs it also appeared on tables whose review reports no harm (the figure is in the eval). It is a prompt to read the study.
- The band has no risk-of-bias assessment; it is our rubric over the abstracts found, not a GRADE rating.
- A meta-analysis and trials it already pooled can sit in the same table.
Models and licences
Standard tier. Licence posture: permissive (Apache, MIT or BSD).
- Evidence engine: PubMed search and fetch, quote and n checks in code, the wording guard, the rubric grade and the signed record (CPU)decosa-evidence engine (decosa_api/studies) with the tool's routes (decosa_api/verticals/studies)Apache-2.0 (engine); AGPL-3.0-or-later (routes)
- Model: reads each abstract (design, n, population, direction, finding, quote), reads it a second time looking only for harm, then judges each finding against the same abstractQwen3.8-27B (NVFP4)Apache-2.0
- Ranks the PubMed results by relevance to the supplement and outcome before any abstract is readQwen3-Reranker-4B (evidence retrieval block)Apache-2.0
All quality evidence
Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.
Lite · one 32 GB card, no reranker (1)
- Table quality without the reranker: not measured yetno run without the reranker has been scored
Standard · one 96 GB card (measured; hosted demo) (5)
- Held-out: the table's verdict matches the systematic review's (benefit, harm or no claim): 231 / 289 (80%)decosa-api docs/evals/what-studies-found.md, 2026-09-30; held-out split, 289 pairs, scored once
- Held-out: calls a benefit where the review found no clear difference: 1 / 61 (2%)decosa-api docs/evals/what-studies-found.md, 2026-09-30; held-out split, 289 pairs, scored once
- Held-out: harms surfaced by the verdict or the harm flag: 36 / 40 (90%)decosa-api docs/evals/what-studies-found.md, 2026-09-30; held-out split, 289 pairs, scored once
- Held-out: harm verdict or flag where the review reports no harm: 46 / 249 (18%)decosa-api docs/evals/what-studies-found.md, 2026-09-30; held-out split, 289 pairs, scored once
- Held-out: finds the benefit where the review found one: 56 / 82 (68%)decosa-api docs/evals/what-studies-found.md, 2026-09-30; held-out split, 289 pairs, scored once