{"schema_version":"1","site":"https://decosa.ai","id":"what-studies-found","num":"183","name":"What studies found","tool_name":"See what studies found for a supplement","short":"What studies found","blurb":"Type a supplement and an outcome, such as probiotics and antibiotic-associated diarrhea. It searches PubMed for trials, meta-analyses and systematic reviews and gives you a table of what they found: one row per study with its design, how many people, who they were, whether the result was better, no different or worse, the finding in numbers where the abstract gives them, and a short quote. Each row links to its PubMed record, each quote is found word for word in the abstract by code, and each finding is judged against its own abstract. A published rubric in code then grades the body of evidence, a separate reading of every abstract flags possible harm, and the studies it left out are listed with why. It reports what studies found: no doses, no advice on what to take.","status":"preview","labels":{"industry":["science-research","healthcare"],"job":["review"],"input":["text"],"deploy":["hosted","selfhost"],"status":"preview","output":["data","record"],"data":["none"],"hardware":"gpu-96","licence":"permissive"},"industries":["science-research","healthcare"],"runs_in":["hosted","selfhost"],"part_of":[],"built_from":["pubmed-evidence","evidence-retrieval","grounding","signed-record"],"models":"Qwen3.8-27B (reads each abstract, reads it again for harm, then checks each finding against it) · Qwen3-Reranker-4B (picks the studies, evidence retrieval block)","where":"Hosted or self-host","hardware":"1x RTX PRO 6000 (96 GB) or 1x RTX 5090 (32 GB) for Qwen3.8-27B; about 8 GB for the reranker; the search, checks and grade run on CPU","final_artifact":"An evidence table (Markdown or CSV) with a PubMed link per row, the grade with the rubric's reasons, the studies left out with why, and a signed record verifiable at /record/verify.","self_host_first":false,"verification":{"receipt_coverage":"full","summary":"Receipt per model call; quotes checked word for word in code; the rubric grades in code; signed hash-chained record of the search, each study and the grade","manual_qa":{"hosted":{"date":"2026-09-30","result":"pass","p50_ms":26328,"p95_ms":31027,"runs":5,"receipts_per_run":32,"cost_per_run_usd":0.022935},"selfhost":null,"known_limits":["Hosted: measured on production on 30 Sep 2026 with the first sample (probiotics and antibiotic-associated diarrhea) through the API, one run at a time. Well-studied pairs like the samples read more reviews and cost more than the average table.","Reads abstracts only; results reported only in the full paper are missed.","The one-line verdict is not a systematic review. On held-out pairs it matches the published review's more often than not but not always (the figures are in the eval); the misses are mostly tables that come out mixed or unclear where the review found a benefit.","The rubric puts the newest Cochrane review first even when it studied a narrower group than the question (the melatonin sample: a review of shift workers decides, the other reviews disagree, and the verdict is \"Mixed results\"). Choosing the review by who it studied is planned, not built.","The same search can give a different verdict on another run: the model's calls on what is relevant vary a little, and PubMed changes.","The harm flag errs toward flagging: on held-out pairs it also appeared on tables whose review reports no harm (the figure is in the eval). It is a prompt to read the study.","The band has no risk-of-bias assessment; it is our rubric over the abstracts found, not a GRADE rating.","A meta-analysis and trials it already pooled can sit in the same table."],"nightly_covers":null},"nightly":"https://api.decosa.ai/verify/status"},"eval_summary":{"metrics":[{"name":"Verdict matches the systematic review's (benefit, harm or no claim)","value":"231 / 289 (80%)","unit":null,"n":289,"split":"heldout","note":"95% CI 75 to 84%. The engine it replaced, on the same pairs: 189 / 289."},{"name":"Calls a benefit where the review found no clear difference","value":"1 / 61 (2%)","unit":null,"n":61,"split":"heldout","note":"The engine it replaced, on the same pairs: 8 / 61."},{"name":"Harms surfaced: the verdict is Worsens, or the table carries its harm flag","value":"36 / 40 (90%)","unit":null,"n":40,"split":"heldout","note":"By the verdict alone: 28 / 40. The engine it replaced, on the same pairs: 18 / 40."},{"name":"Harm verdict or harm flag where the review reports no harm (the price of the flag)","value":"46 / 249 (18%)","unit":null,"n":249,"split":"heldout","note":"A harm verdict alone: 6 / 249. The flag is a prompt to read a study, so it errs toward flagging."},{"name":"Finds the benefit where the review found one","value":"56 / 82 (68%)","unit":null,"n":82,"split":"heldout","note":"The engine it replaced, on the same pairs: 56 / 82. The misses are mostly tables that come out mixed or unclear."},{"name":"Verdict matches exactly, five ways (improves, worsens, no clear difference, mixed, unclear)","value":"180 / 282 (64%)","unit":null,"n":282,"split":"heldout","note":"Pairs where both labellers gave the same five-way label. The engine it replaced, on the same pairs: 135 / 282."},{"name":"The review behind the label was among the studies read","value":"256 / 290 (88%)","unit":null,"n":290,"split":"heldout","note":"Cochrane reviews: 143 / 154. The engine it replaced, on the same pairs: 161 / 290."},{"name":"One abstract read alone: verdict matches the label (benefit, harm or no claim)","value":"290 / 331 (88%)","unit":null,"n":331,"split":"heldout","note":"Harms surfaced per study: 60 / 62."}],"dataset":"290 supplement and outcome pairs, each with a published systematic review, none used while building the engine (83 where the review found a benefit, 40 a harm, 61 no clear difference, 79 too uncertain to tell, 27 mixed). Labels: two blind AI labellers (AI sub-agents) read each review's abstract; a pair counts only when both call it a relevant supplement review and agree on benefit / harm / no claim. Each table was limited to studies published up to the review's year. Scored once, by a rule written down before any output existed; 1 pair could not be run.","held_out":true,"caveats":["The one-line verdict is not a substitute for a systematic review: read the rows and the linked papers.","Labels are by two blind AI sub-agents reading each review's abstract, not by clinicians or systematic reviewers; pairs where they disagreed on benefit, harm or no claim were left out.","Each table was limited to studies published up to the review's year, so the review itself could be found; a search today can read newer studies and say something else.","The harm pairs came from harm topics written down before any search; they are not a random sample of supplement questions.","Abstracts only; no risk-of-bias assessment.","Model calls used the direct route to the same weights as the gateway; cost is at list price from token counts."],"date":"2026-09-30","doc_url":"https://decosa.ai/metrics/evals/what-studies-found"},"stack":{"summary":"Type a supplement and an outcome. It searches PubMed for trials, meta-analyses and systematic reviews, reads each abstract for design, size, population, what was found and a short quote, and checks each finding against its own abstract. A published rubric in code then grades the body of evidence by design, size and consistency. For evidence-based supplement sites, brand regulatory teams, health journalists and researchers who need a first table they can check, cite and rebuild. It reports what studies found: no advice, no doses.","tagline":"A table of what the trials found for a supplement and an outcome, each row linked to its PubMed record and checked against its abstract.","deployment":"hosted-or-self-host","regulatory_note":"Written 29 Sep 2026. This tool reports what published studies found. It is not medical advice, not a recommendation and not a dose, and it says nothing about any product. For supplement sellers: under the Dietary Supplement Health and Education Act of 1994 (DSHEA), a supplement label may carry a structure/function statement with the FDA disclaimer in 21 CFR 101.93(c) (\"This statement has not been evaluated by the Food and Drug Administration. This product is not intended to diagnose, treat, cure, or prevent any disease.\"), but not a disease claim, which 21 CFR 101.93(g) describes as a claim to diagnose, mitigate, treat, cure or prevent disease (https://www.law.cornell.edu/cfr/text/21/101.93, read 29 Sep 2026; official text at https://www.ecfr.gov/current/title-21/chapter-I/subchapter-B/part-101/subpart-F/section-101.93). The FTC's Health Products Compliance Guidance (December 2022, https://www.ftc.gov/business-guidance/resources/health-products-compliance-guidance, read 29 Sep 2026) expects competent and reliable scientific evidence for health claims, which it says will generally need to be randomized, controlled human clinical testing, and the evidence must fit the claim. A table of what trials found is input to that judgement, not substantiation by itself and not wording you can put on a label: the product, amount and population in the trials must match, and counsel decides. The tool's wording guard (studies-guard-1) rewrites or blocks treat, cure, heal, prevent and reverse wording, personal recommendations, and doses or schedules, and a dosing question is refused. The search words go to NCBI's public E-utilities, whose usage policy applies [not re-read for this note]. PubMed abstracts can be under publisher copyright [unverified: it varies by journal], so the tool keeps PMIDs, its own structured findings, quotes of at most 15 words and a SHA-256 of each abstract, not abstract text. Model licences: Apache-2.0 (Qwen3.8-27B, Qwen3-Reranker-4B). Not legal or regulatory advice.","components":[{"id":"engine","role":"Evidence engine: PubMed search and fetch, quote and n checks in code, the wording guard, the rubric grade and the signed record (CPU)","name":"decosa-evidence engine (decosa_api/studies) with the tool's routes (decosa_api/verticals/studies)","hf_repo":null,"license":"Apache-2.0 (engine); AGPL-3.0-or-later (routes)","params":null,"quant":null,"vram_gb":0,"memory_gb_estimate":null,"engine":"Python 3.12; NCBI E-utilities (esearch, efetch); the grounding block (use case 17) judges each finding against its abstract","receipt_coverage":"partial","in_hosted_demo":null,"tiers":["lite","standard"],"alternative_to":null},{"id":"llm","role":"Model: reads each abstract (design, n, population, direction, finding, quote), reads it a second time looking only for harm, then judges each finding against the same abstract","name":"Qwen3.8-27B (NVFP4)","hf_repo":"nvidia/Qwen3.8-27B-NVFP4","license":"Apache-2.0","params":"27.8B","quant":"NVFP4 (MLP NVFP4, GDN/attention FP8) + FP8 KV cache; MTP head, 3 draft tokens","vram_gb":20,"memory_gb_estimate":null,"engine":"vLLM 0.29.0, temperature 0, thinking off","receipt_coverage":"strong","in_hosted_demo":true,"tiers":["lite","standard"],"alternative_to":null},{"id":"reranker","role":"Ranks the PubMed results by relevance to the supplement and outcome before any abstract is read","name":"Qwen3-Reranker-4B (evidence retrieval block)","hf_repo":"Qwen/Qwen3-Reranker-4B","license":"Apache-2.0","params":"4B","quant":"BF16","vram_gb":null,"memory_gb_estimate":8.1,"engine":"transformers (services/retrieval)","receipt_coverage":"partial","in_hosted_demo":true,"tiers":["standard"],"alternative_to":null}],"tiers":[{"id":"lite","label":"Lite · one 32 GB card, no reranker","summary":"The same model and checks on one smaller card, with PubMed's own relevance order instead of the reranker. Fewer of the picked studies may be on point.","components":["engine","llm"],"hardware":"1x RTX 5090 32 GB (estimate)","quality_evidence":[{"metric":"Table quality without the reranker","value":"not measured yet","source":"no run without the reranker has been scored"}],"latency_note":"estimate: similar to standard; not measured on this card.","in_hosted_demo":false,"receipt_coverage":"partial","receipt_note":"Self-hosted: your own box signs each model call and the record.","hosting":null},{"id":"standard","label":"Standard · one 96 GB card (measured; hosted demo)","summary":"Qwen3.8-27B reads every abstract, reads it again for harm and checks each finding; the reranker picks the studies; the rubric grades in code.","components":["engine","llm","reranker"],"hardware":"1x RTX PRO 6000 96 GB, plus about 8 GB for the reranker","quality_evidence":[{"metric":"Held-out: the table's verdict matches the systematic review's (benefit, harm or no claim)","value":"231 / 289 (80%)","source":"decosa-api docs/evals/what-studies-found.md, 2026-09-30; held-out split, 289 pairs, scored once"},{"metric":"Held-out: calls a benefit where the review found no clear difference","value":"1 / 61 (2%)","source":"decosa-api docs/evals/what-studies-found.md, 2026-09-30; held-out split, 289 pairs, scored once"},{"metric":"Held-out: harms surfaced by the verdict or the harm flag","value":"36 / 40 (90%)","source":"decosa-api docs/evals/what-studies-found.md, 2026-09-30; held-out split, 289 pairs, scored once"},{"metric":"Held-out: harm verdict or flag where the review reports no harm","value":"46 / 249 (18%)","source":"decosa-api docs/evals/what-studies-found.md, 2026-09-30; held-out split, 289 pairs, scored once"},{"metric":"Held-out: finds the benefit where the review found one","value":"56 / 82 (68%)","source":"decosa-api docs/evals/what-studies-found.md, 2026-09-30; held-out split, 289 pairs, scored once"}],"latency_note":"measured on our server 2026-09-30, 3 tables at once, live PubMed: p50 29.9 s and p95 55.9 s per table; $0.0142 per table on average at list price (21.3 model calls). The 7 recorded sample runs took 34.47 to 72.79 s and cost $0.018512 to $0.022449 each.","in_hosted_demo":true,"receipt_coverage":"strong","receipt_note":"Hosted: each extraction, harm reading and grounding call has a gateway-signed receipt; the rerank has a model-call receipt; the record lists them.","hosting":null}],"alternates":[],"services":[{"name":"decosa-api","port":8000,"image":"${DECOSA_REGISTRY}/decosa-api:<tag>","purpose":"GET /evidence/table/info and /evidence/table/samples; POST /evidence/table (SSE or JSON). Keeps the PMIDs each search returned (keyed by the search term's SHA-256, 7 days); titles and abstracts live in memory for the request."},{"name":"vLLM","port":8114,"image":"vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1","purpose":"Qwen3.8-27B NVFP4 behind our gateway (hosted) or called directly (self-host)."},{"name":"decosa-retrieval (optional)","port":8499,"image":"built from decosa-api services/retrieval/Dockerfile (RETRIEVAL_RERANK=qwen3-rr-4b)","purpose":"The reranker; set DECOSA_RETRIEVAL_URL. With DECOSA_STUDIES_RERANK=off the engine uses PubMed's order."}],"tools":[{"name":"NCBI E-utilities (PubMed)","url":"https://www.ncbi.nlm.nih.gov/books/NBK25501/","license":"Public service of the US National Library of Medicine; PubMed records are public, abstracts may be under publisher copyright","purpose":"esearch for the PMIDs (trials, meta-analyses and systematic reviews; retractions, errata, comments, editorials and letters excluded), efetch for titles and abstracts. Identifies itself with a tool name and, if set, a contact address and API key."},{"name":"Grounding block (tool 17)","url":null,"license":"AGPL-3.0-or-later","purpose":"Judges each row's finding against its own abstract: supported, partial, unsupported or contradicted. Unsupported and contradicted rows are dropped before grading and listed as left out."},{"name":"Rubric studies-rubric-4 and rule studies-net-2 (decosa_api/studies/grade.py, direction.py, safety.py)","url":null,"license":"Apache-2.0","purpose":"Works out each study's direction in code from separate fields (the measure's polarity, the raw change, significance, the authors' conclusion and stated certainty), then grades the body: the best available review decides (the newest Cochrane review first), its stated GRADE certainty is the band, reviews that disagree give \"Mixed results\", and trials alone need minimum sizes. \"No clear difference\" and \"Unclear: evidence too weak to tell\" are two different verdicts. Harm is read separately: when any study read reports a worse result for the outcome, the table carries a harm flag (\"Possible harm: check the studies\") whatever the verdict. GET /evidence/table/info returns the full text."},{"name":"POST /record/verify","url":null,"license":"AGPL-3.0-or-later","purpose":"Checks the signed record and names the first entry that was changed."}],"hardware":[{"tier":"1x RTX PRO 6000 Blackwell 96 GB","fits":true,"notes":"Measured on our server: the eval and the recorded demo runs ran on this card, shared with other services, with the reranker on a second card."},{"tier":"1x RTX 5090 32 GB","fits":true,"notes":"Estimate: Qwen3.8-27B NVFP4 needs about 20 GB of weights plus KV cache for abstracts of a few thousand tokens; run without the reranker or put it on a second card. Not run here on a 5090."}],"latency":[{"lane":"One table (8 studies asked for, up to 5 reviews always read), held-out split, 3 tables at once, live PubMed","typical_ms":29900,"source":"measured on our server 2026-09-30: p50 29.9 s, p95 55.9 s over 289 tables (decosa-api docs/evals/what-studies-found.md)"}],"benchmark":null,"notes":["Reads abstracts only: a result reported only in the full paper is missed, and so are secondary outcomes.","A meta-analysis and the trials it pooled can both be rows in the same table; the rubric weighs syntheses more but does not remove the overlap yet.","The band is GRADE-inspired, not GRADE: it has no risk-of-bias assessment, so without a review's stated certainty it stops at moderate.","The rubric puts the newest Cochrane review first even when it studied a narrower group than you asked about (shift workers, children, pregnancy). When reviews of different groups disagree the verdict is \"Mixed results\", and the result shows who each review studied.","The harm flag errs toward flagging: it is a prompt to read a study, not a finding of harm.","You choose how many studies to read (the API's limits are in GET /evidence/table/info); reviews are always read, so a table can hold a few more studies than you asked for. Ask for primary trials only with designs: \"primary\", or limit to studies published before a year with published_before."]},"buyer_facts":[{"label":"Data retention","value":"The hosted service keeps the PMIDs each search returned for 7 days, keyed by a SHA-256 of the search term, so repeat searches are fast. Titles and abstracts live in memory for the request; the signed record holds each abstract's SHA-256, not its text."},{"label":"What leaves the box","value":"The search words and PMIDs go to NCBI's public PubMed service. Abstracts and search words go to the model and the reranker, which run on Decosa's hosted service (hosted) or your machine (self-host)."},{"label":"What it will not do","value":"No doses, no advice on what to take, no treat, cure or prevent wording: a dosing or advice question is refused, and the guard rewrites such wording in findings."},{"label":"Input","value":"A supplement (up to 80 characters) and an outcome (up to 120), or a sample. Optional: how many studies to read, primary trials only, or only studies published before a year."},{"label":"Output","value":"A table with one row per study (PMID and link, design, n, population, direction, finding, a quote of at most 15 words, the grounding verdict, a harm flag), the grade with its reasons, the table's harm flag with the studies behind it, the studies left out and why, and a signed record verifiable at /record/verify."}],"data_handling":{"page":"/data#what-studies-found","self_host":{"level":"confidential","leaves":"content","summary":"Runs on your machine, but by default sends content to the services listed in external_calls."},"hosted":{"level":"operator-processed","demo_only":false,"summary":"TLS to Decosa's server, then decrypted and processed by Decosa's API server, with the open models run by NEAR AI through OpenRouter, with Reka AI as the only fallback under Decosa's account.","gpus":"operator-contracted","third_parties":["NCBI E-utilities"],"retention":"The hosted service keeps the PMIDs each search returned for 7 days, keyed by a SHA-256 of the search term, so repeat searches are fast. Titles and abstracts live in memory for the request; the signed record holds each abstract's SHA-256, not its text.","used_for_training":false,"encrypted_while_processed":false},"sealed_tier":{"applies":false,"note":"The sealed tier (raw chat only, never use-case pipelines) is paused at launch (/docs/sealed-tier)."},"external_calls":[{"to":"NCBI E-utilities (PubMed, eutils.ncbi.nlm.nih.gov)","route":"both","sends":"content","what":"The search words (the supplement and the outcome, in a PubMed query) and the PMIDs to fetch. PubMed returns public records: titles and abstracts.","default":"always","off":null}]},"console":{"href":"/tools/life-sciences/what-studies-found","input":"runner","lanes":[{"id":"table","title":"What studies found","kind":"list"},{"id":"grade","title":"Grade and studies left out","kind":"markdown"},{"id":"record","title":"Signed record","kind":"json"}],"samples":[{"n":1,"id":"probiotics-antibiotic-diarrhea","title":"Probiotics antibiotic diarrhea","deep_link":"/tools/life-sciences/what-studies-found?sample=1&autorun=0"},{"n":2,"id":"melatonin-sleep-onset","title":"Melatonin sleep onset","deep_link":"/tools/life-sciences/what-studies-found?sample=2&autorun=0"},{"n":3,"id":"beta-carotene-lung-cancer","title":"Beta carotene lung cancer","deep_link":"/tools/life-sciences/what-studies-found?sample=3&autorun=0"},{"n":4,"id":"caffeine-blood-pressure","title":"Caffeine blood pressure","deep_link":"/tools/life-sciences/what-studies-found?sample=4&autorun=0"},{"n":5,"id":"vitamin-d-asthma","title":"Vitamin d asthma","deep_link":"/tools/life-sciences/what-studies-found?sample=5&autorun=0"},{"n":6,"id":"omega3-depression","title":"Omega3 depression","deep_link":"/tools/life-sciences/what-studies-found?sample=6&autorun=0"},{"n":7,"id":"cranberry-uti","title":"Cranberry uti","deep_link":"/tools/life-sciences/what-studies-found?sample=7&autorun=0"}],"deep_link_params":{"sample":"1-based index into samples, or a sample id","autorun":"1 = start the run once the sample is loaded; 0 (default) = only preselect","reduce-motion":"1 = turn off animations"}},"api":{"base":"https://api.decosa.ai","contract":"/api/contract.json","contract_markdown":"/api/contract.md","reference":"/docs/api","keys":"/account/keys"},"prompts":{"hosted":"/prompts/what-studies-found-hosted.md","selfhost":"/prompts/what-studies-found-selfhost.md","assemble":"/prompts/what-studies-found-assemble.md","mac":null},"rehearsal":{"bundle":"/samples/what-studies-found.zip","bundle_url":"https://decosa.ai/samples/what-studies-found.zip","folder":"/samples/what-studies-found/","expected":"/samples/what-studies-found/expected.json","files":["/samples/what-studies-found/expected.json"],"bytes":1457,"checks":["probiotics and antibiotic-associated diarrhea: the studies found a better result with probiotics","the verdict is a benefit claim, labelled Improves","a systematic review decides the verdict","the band is low or better (it is the certainty the best review states)","at least four studies are counted","the table carries a harm flag (none, possible or reported) computed from every study read","the rubric in force is the one that reads harm asymmetrically","every row links to PubMed","the record is signed and verifies","melatonin and sleep onset: the reviews read disagree, so the verdict is Mixed results","a mixed verdict claims neither a benefit nor a harm","the rubric's reasons say the reviews disagree","the table shows why: one review read is about shift workers, the others are not","a dose request is refused"],"licence":"No input files: the search runs live against public PubMed records (NCBI E-utilities). Abstracts are read for the request and not kept.","about":"Builds the evidence table for probiotics and antibiotic-associated diarrhea from PubMed (a well-studied pair: a Cochrane review and several meta-analyses of randomised trials found fewer cases with probiotics), then the table for melatonin and sleep onset latency, where the reviews disagree (the Cochrane review read is about shift workers, the other reviews are about people with sleep problems) and the verdict is \"Mixed results\", then asks for a dose and must be refused. The tables must link every row to PubMed by PMID, keep quotes at 15 words or fewer, grade the body with the published rubric, carry the harm flag, and end in a signed record that verifies.","run":{"containers":"docker compose exec api python scripts/rehearse.py what-studies-found","checkout":"python scripts/rehearse.py what-studies-found --bundle what-studies-found.zip --base-url http://127.0.0.1:8445","mac":".venv/bin/python scripts/rehearse.py what-studies-found"},"guidance":"Set up with a coding agent (we recommend Claude Code with Claude Opus 5.5; any capable coding agent works) on mock data only, run the rehearsal until every check passes, then run your own data locally yourself. Never give the agent real data during setup."},"hardware_fit":{"check":"/self-host/hardware?use=what-studies-found","data":"/api/hardware.json","tiers":[{"id":"lite","gpu_gb":57.6,"basis":"stack","unknown":[]},{"id":"standard","gpu_gb":68.2,"basis":"estimate","unknown":[]}],"mac":null},"links":{"page":"/tools/life-sciences/what-studies-found","json":"/use-cases/what-studies-found.json","metrics":"/metrics/what-studies-found","console":"/tools/life-sciences/what-studies-found","console_sample":"/tools/life-sciences/what-studies-found?sample=1&autorun=0","stack":"/tools/life-sciences/what-studies-found#stack","try_live":"/tools/life-sciences/what-studies-found","watch":"/tools/life-sciences/what-studies-found","build":"/tools/life-sciences/what-studies-found#build","self_host":"/tools/life-sciences/what-studies-found#self-host","prompts":{"hosted":"/prompts/what-studies-found-hosted.md","selfhost":"/prompts/what-studies-found-selfhost.md","assemble":"/prompts/what-studies-found-assemble.md","mac":null}}}