Eval: See what studies found for a supplement (use case 183, what-studies-found)
29 Sep 2026. Engine decosa_api/studies, rubric studies-rubric-2, Qwen3.8-27B. The run used our server's direct route, which serves the same weights as the gateway. Cost is at list price ($0.30/M input, $1.50/M output) from token counts. Wall time is measured on a pre-release server under load (3 tables at once), live PubMed included.
What is measured
- Verdict against Cochrane. For each supplement-outcome pair where a Cochrane review states a conclusion, we check whether the table's body direction and band match the review's direction (improved / no difference / worsened / uncertain) and its GRADE certainty.
- Each table only sees studies published up to the review's year.
- Two modes: A, as deployed (trials plus their meta-analyses and reviews, so the Cochrane review itself can be among the rows), and B, primary trials only.
- Every row against its own source. The grounding block (use case 17) judges every row's finding against that study's abstract automatically. Code checks every quote word for word. A blind judge checks a sample of rows field by field.
- Cost and time per table.
Labels came from blind sub-agents (Claude Code Opus 5.5) that read only the Cochrane abstracts and never saw engine output: evaldata/what-studies-found/set1.json and set2-heldout.json. They hold the review PMID, direction and certainty; abstract text is left out. There are two sets:
- Set 1: 40 pairs.
- Set 2: 31 held-out pairs from different reviews, made after rubric 2 was written.
Runner: evaldata/what-studies-found/eval_run.py, then summarize.py.
Tuning disclosure
Rubric 1 overstated certainty on set 1: 11 of the 24 pairs Cochrane rated low or very low came out "high". After seeing that, we wrote rubric 2:
- the newest agreeing review's stated GRADE certainty caps the band;
- without one, the band is at most "moderate", because abstracts can't show risk of bias;
- a "very low" band states no direction.
Rubric 2 also extracts stated_certainty and drops joining words from PubMed terms. One set-1 pair returned 0 hits for "respiratory infections in young children", and 101 without "in".
Set 1 is therefore seen data for rubric 2. Set 2 is the honest number.
Results
| Set | Mode | Pairs | Direction agrees | Band exact | Band within 1 | Rubric 1 on the same rows (dir / exact / within 1) |
|---|---|---|---|---|---|---|
| Set 2 (held-out) | A | 31 | 16 (52%) | 14 (45%) | 26 (84%) | 16 / 7 / 18 |
| Set 2 (held-out) | B | 31 | 18 (58%) | 10 (32%) | 24 (77%) | 16 / 7 / 22 |
| Set 1 (seen) | A | 40 | 26 (65%) | 20 (50%) | 33 (83%) | 23 / 10 / 24 |
| Set 1 (seen) | B | 40 | 15 (38%) | 16 (40%) | 31 (78%) | 14 / 15 / 30 |
- Where direction disagrees (set 2, A): 4 no_difference→improved, 3 uncertain→improved, 2 improved→no_difference and 2 uncertain→no_difference. The typical miss is a review that concludes "little or no difference" or "uncertain" while the abstracts the search returns read positive: small positive trials and older, optimistic meta-analyses.
- Opposite directions (improved against no_difference or worsened): 6 of 31 (A) and 1 of 31 (B).
- Rubric 2 against rubric 1: band agreement doubles on the held-out set (exact 7→14, within one 18→26). Direction is unchanged.
Rows against their sources
- Grounding block (every row): set 2 A, 118 supported and 10 partial of 128 rows; set 1 A, 188 supported and 10 partial of 198. Rows judged unsupported or contradicted are dropped before grading and listed as excluded.
- Quotes found word for word by code: 116 of 128 (set 2 A). A quote that fails the check is removed, so no row shows an unverified quote. Quotes are at most 15 words.
- Blind judge (Claude Code Opus 5.5 sub-agent). It checked 100 rows sampled from set 1, stratified by direction, and saw only the abstract and the row; it never saw labels or grounding verdicts.
| Field | Result (of 100) | Main errors |
|---|---|---|
| Relevant to the question | 92 | Dietary-intake or observational meta-analyses; mixed products; the supplement in both arms |
| Design | 88 correct, 8 unclear, 4 wrong | Mostly "meta-analysis" for a review that didn't pool |
| n | 82 correct, 18 not stated, 0 wrong | |
| Direction | 96 correct | |
| Finding | 93 supported, 7 partly, 0 unsupported or contradicted | |
| Quote | 99 verbatim, 1 empty |
- Grounding against the judge: the grounding block flagged 3 of the judge's 7 "partly" rows as partial. It also flagged 6 rows as partial that the judge found fully supported, so it is conservative.
Cost and time per table (8 studies, 17–20 model calls)
- Set 2 A: p50 $0.0068 and 8.7 s; p95 $0.0078 and 11.4 s.
- Set 1 A: p50 $0.0071 and 9.1 s; p95 $0.0084 and 12.0 s.
That is list price, direct route. Gateway runs are metered at the same price.
What this means
- The table is accurate enough to publish with review: rows cite real PMIDs, directions are right 96% of the time, findings hold up against the abstract, and quotes are verbatim.
- The one-line verdict is not a substitute for a systematic review. On held-out pairs it agrees with Cochrane's direction about half the time, and it tends to read small positive trials as a benefit.
- Products should lead with the table, show the band as "our rubric over the abstracts found", and link the rubric. Dr. Grey's proposals require human review before apply.
Checkable properties of the sample run (rehearsal/what-studies-found)
- Probiotics and antibiotic-associated diarrhea: the direction is
improved, the labelImproves, decided by a systematic review. - The band is
high,moderateorlow(the certainty the best review states). - At least 4 studies are counted, and the table carries a harm flag (
none,possibleorreported). - Every row links to https://pubmed.ncbi.nlm.nih.gov/<PMID>/.
- The signed record verifies at /record/verify.
- Melatonin and sleep onset latency: the label is
Mixed results, the claimnone, the reasons say the reviews disagree, and a row's population is shift workers (the Cochrane review that takes best-review standing). - A dosing question is refused with 400 ("no dosing").
What would move the numbers
- Full text from PMC Open Access (CC-BY), so results reported only in the paper are read.
- Deduplicating trials that a meta-analysis in the same table already pools.
- A review-first mode that leads with the newest systematic review's conclusion.
- Our own extractor (M-ev, page 82) trained on PMC-OA abstracts.
Harm read asymmetrically, wider review search, fresh split C (30 Sep 2026, the pre-release branch, frozen at b9108c64)
Prompts studies-extract-3 + studies-harm-1, net rule studies-net-2, rubric studies-rubric-4, against the engine that was
live (6cabc7de: prompt 1, rubric 2). This supersedes the "Direction and meaning" section below: that version fixed the false
benefits but lost harms (per study 26/27 -> 22/27 on its held-out split), so it was not merged. What changed since:
- Harm is its own pass. Every abstract gets a second model call that looks only for harm: a worse result for the outcome
(reported / possible / none, with the sentence, checked word for word in the abstract) and what the abstract says about
adverse events in general. A reported harm makes the row "worsens" whatever the exposure type (a trial, a review, an
association with intake or blood levels); it never overrules a statistic of the main reading, it flags the row instead. The
table carries
harm_signal("Possible harm: check the studies") when any study read reports a worse result. - Add-on comparators count. "Chemotherapy alone", "vitamin D alone": both groups got the same treatment and one also got the supplement. Only a head-to-head comparison with a different treatment is set aside.
- Review search. Supplement name variants and synonyms, the outcome with and without its condition, its other names from the polarity resource, reviews with the supplement in the title, the newest Cochrane reviews; up to 5 reviews always read (3 Cochrane, the newest version of each), added on top of the studies asked for.
Fresh held-out split C (never used for tuning; scored once)
Body of evidence (one table per review; the table's verdict against the review's), n = 289
| Metric | Old engine (live) | New engine |
|---|---|---|
| Verdict, three-way (benefit / harm / no claim): MERGE RULE | 189/289 (65%, 95% CI 60 to 71%) | 231/289 (80%, 95% CI 75 to 84%) |
| False benefit on null reviews (label: no clear difference): MERGE RULE | 8/61 (13%, 95% CI 7 to 24%) | 1/61 (2%, 95% CI 0 to 9%) |
| Harm surfaced (verdict Worsens, or the harm flag): MERGE RULE | 18/40 (45%, 95% CI 31 to 60%) | 36/40 (90%, 95% CI 77 to 96%) |
| Harm recall, strict (verdict Worsens only) | 18/40 (45%, 95% CI 31 to 60%) | 28/40 (70%, 95% CI 55 to 82%) |
| False harm verdict (label is not a harm, verdict Worsens): guard | 10/249 (4%, 95% CI 2 to 7%) | 6/249 (2%, 95% CI 1 to 5%) |
| Harm verdict or flag on a label that is not a harm (the price of the flag) | 10/249 (4%, 95% CI 2 to 7%) | 46/249 (18%, 95% CI 14 to 24%) |
| False benefit, null or too-weak-to-tell reviews | 36/140 (26%, 95% CI 19 to 34%) | 11/140 (8%, 95% CI 4 to 14%) |
| Benefit recall | 56/82 (68%, 95% CI 58 to 77%) | 56/82 (68%, 95% CI 58 to 77%) |
| Verdict, five-way exact (both labellers agree on the five-way label) | 135/282 (48%, 95% CI 42 to 54%) | 180/282 (64%, 95% CI 58 to 69%) |
| Review found: the labelled review among the studies read | 161/290 (56%, 95% CI 50 to 61%) | 256/290 (88%, 95% CI 84 to 91%) |
| the labelled review among the counted rows | 129/290 (44%, 95% CI 39 to 50%) | 205/290 (71%, 95% CI 65 to 76%) |
| Cochrane pairs: the labelled Cochrane review among the studies read | 87/154 (56%, 95% CI 49 to 64%) | 143/154 (93%, 95% CI 88 to 96%) |
| Cochrane pairs: among the counted rows | 64/154 (42%, 95% CI 34 to 49%) | 113/154 (73%, 95% CI 66 to 80%) |
| "Too few studies to tell" | 30/290 (10%, 95% CI 7 to 14%) | 22/290 (8%, 95% CI 5 to 11%) |
Per study (the engine reads the review's own abstract), n = 331
| Metric | Old engine (live) | New engine |
|---|---|---|
| Verdict, three-way (benefit / harm / no claim) | 280/331 (85%, 95% CI 80 to 88%) | 290/331 (88%, 95% CI 84 to 91%) |
| False benefit on null reviews (label: no clear difference) | 1/70 (1%, 95% CI 0 to 8%) | 0/70 (0%, 95% CI 0 to 5%) |
| Harm surfaced (verdict Worsens, or the harm flag) | 44/62 (71%, 95% CI 59 to 81%) | 60/62 (97%, 95% CI 89 to 99%) |
| Harm recall, strict (verdict Worsens only) | 44/62 (71%, 95% CI 59 to 81%) | 51/62 (82%, 95% CI 71 to 90%) |
| False harm verdict (label is not a harm, verdict Worsens): guard | 3/269 (1%, 95% CI 0 to 3%) | 8/269 (3%, 95% CI 2 to 6%) |
| Harm verdict or flag on a label that is not a harm (the price of the flag) | 3/269 (1%, 95% CI 0 to 3%) | 38/269 (14%, 95% CI 10 to 19%) |
| False benefit, null or too-weak-to-tell reviews | 12/153 (8%, 95% CI 5 to 13%) | 1/153 (1%, 95% CI 0 to 4%) |
| Benefit recall | 69/84 (82%, 95% CI 73 to 89%) | 68/84 (81%, 95% CI 71 to 88%) |
| Verdict, five-way exact (both labellers agree on the five-way label) | 229/324 (71%, 95% CI 66 to 75%) | 258/324 (80%, 95% CI 75 to 84%) |
Set: 389 abstracts labelled by both blind labellers; 331 are supplement reviews with a result for the outcome by both (8 more left out because the labellers disagree on benefit / harm / no claim); 290 body pairs. Labels of the body pairs (labeller A): 83 improves, 79 uncertain, 61 no clear difference, 40 worsens, 27 mixed. By source: 151 cochrane-pool, 75 harm-topic, 36 safety-pool, 28 general-pool. The two labellers give the same five-way label in 324/339 (96%, 95% CI 93 to 97%) of the cases both call relevant.
One more pair could not be run for the new engine (PubMed answered HTTP 429 while another of our runs shared the address); it is left out of both columns. Its label is "improves" and the old engine's verdict on it was "unclear", so counting it against the new engine changes no comparison.
Merge rule (set by the lead before the run): three-way verdict, false benefit on null reviews and harm recall each no worse than the old engine, at least two better. Three-way 189/289 -> 231/289; false benefit on null reviews 8/61 -> 1/61; harm surfaced 18/40 -> 36/40 (strict, the verdict itself: 18/40 -> 28/40). Met.
Where the 40 labelled harms went (new engine): 28 have the verdict "Worsens", 8 carry the "possible harm" flag under another verdict, 4 are not surfaced (3 Too few studies to tell; 1 No clear difference; in 3 of them the labelled review was not among the studies read).
The price of the flag: 46/249 tables whose labelled review reports no harm carry a harm verdict or flag (27 harm-topic, 11 cochrane-pool, 6 safety-pool, 2 general-pool: most are harm topics, where other studies read do report a worse result). A false harm VERDICT: 10/249 old, 6/249 new.
Remaining failure types (new engine, body, label -> verdict): improves -> mixed: 13; uncertain -> improves: 10; improves -> unclear: 8; worsens -> unclear: 5; worsens -> no_clear_difference: 5; mixed -> improves: 4; mixed -> worsens: 4; improves -> no_clear_difference: 4; worsens -> mixed: 2; improves -> worsens: 1; no_clear_difference -> improves: 1; no_clear_difference -> worsens: 1. By where the labelled review went: the labelled review counted, another study decided or it was read differently: 25; the labelled review was not read: 13; read, an association with intake or levels (no harm): 9; read, called not about this supplement and outcome: 4; read, set aside as a head-to-head comparison: 3; read, no usable result: 2; read, its finding did not hold up against its abstract: 2. Benefits missed (26 of 82, the same count as the old engine): Mixed results: 13; Too few studies to tell: 5; No clear difference: 4; Unclear: evidence too weak to tell: 3; Worsens: 1. Per study, a false harm verdict rose from 3/269 to 8/269: label mixed, decided by harm reading, exposure given: 1; label mixed, decided by harm reading, exposure intake_or_levels: 3; label mixed, decided by polarity, exposure intake_or_levels: 2; label uncertain, decided by polarity, exposure given: 1; label improves, decided by polarity, exposure intake_or_levels: 1.
Review-found rate: the labelled review is among the studies read in 256/290 tables (old 161/290), and among the counted rows in 205/290 (old 129/290). For pairs whose labelled review is a Cochrane review: read 143/154 (old 87/154), counted 113/154 (old 64/154). In 222 tables a Cochrane candidate was found and its newest version was read; in 67 no Cochrane review matched the search; in 0 one was found and not read.
Method. 389 review abstracts, none used in any earlier split: every unused Cochrane supplement review from 2005 on with a supplement word in the title, the top 2 PubMed reviews for harm topics written down before any search (some are expected nulls), a seeded sample of safety-titled supplement meta-analyses and a seeded sample of general supplement meta-analyses of randomised trials. The supplement and outcome of the unframed cases were set by the model (no verdict asked) before labelling. Two blind Claude Code Opus labellers, in separate folders, each given the same supplement and outcome and neutral case ids; the second never saw the first's fields. A case where they disagree on benefit / harm / no claim is left out. The engine was frozen (b9108c64) and the scoring rule written down before any engine output on this split existed. Both engines ran in-process on our server's direct route (same weights as the gateway), temperature 0, one shared PubMed cache, tables limited to studies published up to the review's year.
Cost and time per table (8 studies asked for, up to 11 read): new 21.3 model calls and $0.0142 at list price (old 10.8 calls, $0.0056); wall time p50 30 s, p95 56 s with 4 tables at a time and another run sharing the model (old p50 15 s).
Dev (the earlier fresh split, read before this work, so tuning data): kept with the eval harness, not reported here as a result.
The samples (30 Sep 2026): chosen for a stable verdict, engine unchanged
The engine merged as frozen. The flagship sample changed: melatonin and sleep onset latency reads "Mixed results" on this engine (a Cochrane review about shift workers takes best-review standing and the general reviews disagree), so the first sample is now a pair from the held-out split that the engine got right, whose verdict was the same in every run below and matches its Cochrane review. Melatonin stays as the second sample and shows the disagreement as it is. Runs had no date limit, as on the site.
| Sample | Supplement and outcome | In process, 4 rounds | Through the API, 4 runs | Deciding review |
|---|---|---|---|---|
probiotics-antibiotic-diarrhea |
probiotics / antibiotic-associated diarrhea | Improves / moderate: 4 of 4 | Improves / moderate: 4 of 4 | PMID 31039287 (Cochrane) |
melatonin-sleep-onset |
melatonin / sleep onset latency | Mixed results / very low: 4 of 4 | Mixed results / very low: 4 of 4 | PMID 25113164 (Cochrane) |
beta-carotene-lung-cancer |
beta-carotene / lung cancer | Worsens / moderate: 4 of 4 | Worsens / moderate: 4 of 4 | PMID 37702300 (Cochrane) |
caffeine-blood-pressure |
caffeine / blood pressure | Mixed results / very low: 4 of 4 | Mixed results / very low: 4 of 4 | PMID 38057002 |
vitamin-d-asthma |
vitamin D / severe asthma exacerbations | No clear difference / high: 4 of 4 | No clear difference / high: 4 of 4 | PMID 36744416 (Cochrane) |
omega3-depression |
omega-3 / depressive symptoms | Unclear: evidence too weak to tell / very low: 4 of 4 | Unclear: evidence too weak to tell / very low: 4 of 4 | PMID 39564892 (Cochrane) |
cranberry-uti |
cranberry / urinary tract infection | Improves / moderate: 4 of 4 | Improves / moderate: 4 of 4 | PMID 37068952 (Cochrane) |
Candidates run and not used (in process, 4 rounds each): iron / anaemia (No clear difference / low: 4 of 4); creatine / muscle hypertrophy (Improves / low: 3 of 4; Mixed results / very low: 1 of 4); St John's wort / depressive symptoms (Improves / moderate: 4 of 4); psyllium / fasting blood sugar (Improves / moderate: 4 of 4); curcumin / fasting blood glucose (Improves / moderate: 4 of 4); iron / restless legs syndrome severity (Improves / moderate: 4 of 4); myo-inositol / gestational diabetes (Improves / low: 4 of 4); soluble fiber / serum lipid profile (Improves / moderate: 4 of 4); zinc / common cold duration (Improves / low: 4 of 4); vitamin C / common cold duration (Improves / moderate: 4 of 4); magnesium / diarrhoea (Worsens / moderate: 4 of 4); vitamin A / hip fracture (Worsens / moderate: 4 of 4); berberine / gastrointestinal adverse events (Mixed results / very low: 4 of 4); omega-3 fatty acids / cognitive function (Mixed results / very low: 4 of 4); vitamin A / all-cause mortality (No clear difference / high: 4 of 4).
The 28 API runs cost $0.0182 to $0.0231 each at list price and took 11.8 to 18.2 s (direct route, one table at a time).
The machine-readable scores for split C (the scorer's own output) are in docs/evals/what-studies-found-split-c.json; the site's figures are generated from it.
Direction and meaning (30 Sep 2026, on a pre-release build, measured at 4f2101b3; superseded by the section above, never merged on its own)
Prompt studies-extract-2, net rule studies-net-1, rubric studies-rubric-3, polarity resource decosa-polarity-1, against the
engine above (prompt 1, rubric 2). Labels by a blind Claude Code Opus sub-agent from the source text, a second blind labeller on
every disagreement (cases where the two differ are left out). Old and new ran on the direct route with one shared PubMed cache.
Tables below are generated from the score files.
Held-out split (never used for tuning)
| Metric | Set | Old | New |
|---|---|---|---|
| Polarity accuracy (outcome names; labelled higher/lower) | 78 names | 61/78 (78%) (Dr. Grey site rules) | 56/78 (72%); 4 wrong, 18 no answer |
| Polarity on names with no good direction (context, neutral, unclear) | 15 names | claims a direction on all | abstains on 10/15 (67%) |
| Polarity accuracy (per study, measure in the abstract) | 67 studies | 31/31 (100%) (implied, only where it can be derived) | 66/67 (99%) |
| Net direction, per study (exact, 5 classes) | 68 abstracts | 57/68 (84%) | 60/68 (88%) |
| Net direction, body of evidence (table verdict vs the review; exact) | 40 pairs | 33/40 (82%) | 29/40 (72%) |
| Same, claim level (benefit / harm / no claim) | 40 pairs | 35/40 (88%) | 37/40 (92%) |
| False-benefit rate on null reviews (review says no clear difference; table says improved) | 18 pairs | 2/18 (11%) | 0/18 (0%) |
| False benefit, null or uncertain reviews | 29 pairs | 3/29 (10%) | 1/29 (3%) |
| False benefit reading the null review's own abstract | 18 abstracts | 0/18 (0%) | 0/18 (0%) |
| Benefit recall, body of evidence | 10 pairs | 9/10 (90%) | 9/10 (90%) |
| Harm recall, per study (label worsens) | 6 abstracts | 5/6 (83%) | 6/6 (100%) |
| Harm precision, per study | 5/5 (100%) | 6/6 (100%) | |
| Harm recall, body of evidence | 1 pairs | 0/1 (0%) | 0/1 (0%) |
| The review itself among the studies read | 40 pairs | 13 | 32 |
Dr. Grey stored rows (old = origin/main outcomeHelpers.ts; new = interpret_row with the polarity resource; labels from the linked abstracts):
| Metric | Rows | Site today | New |
|---|---|---|---|
| Shown as "Worsens" when the evidence is not a harm | 42 | 24/42 (57%) | 2/42 (5%) |
| Verdicts shown (Improves/Worsens) that are wrong | 27/35 (77%) | 4/14 (29%) | |
| Real harms shown as "Worsens" (harm recall) | 4 | 2/4 (50%) | 4/4 (100%) |
| Net matches the label (rows with a readable label) | 37 | 8/37 (22%) | 13/37 (35%) (most rows hold no usable direction, so "unclear" is the honest display) |
| Polarity of the row's outcome name | 42 | 42/42 (100%) | 40/42 (95%) (2 no answer) |
The same rows answered by the engine (one table per row, supplement + the row's outcome name, no date limit):
| Metric | Rows | Old engine | New engine |
|---|---|---|---|
| Net matches the label (exact) | 37 | 20/37 (54%) | 13/37 (35%) |
| Claim level | 37 | 24/37 (65%) | 17/37 (46%) |
| False benefit (label no clear difference or mixed) | 12 | 4/12 (33%) | 3/12 (25%) |
| Benefit recall | 21 | 15/21 (71%) | 9/21 (43%) |
| Harm recall | 4 | 1/4 (25%) | 0/4 (0%) |
Contested labels left out: 2 study cases, 0 pairs, 5 rows, 1 names.
Dev split (rules and prompt were tuned on these)
| Metric | Set | Old | New |
|---|---|---|---|
| Polarity accuracy (outcome names; labelled higher/lower) | 98 names | 78/98 (80%) (Dr. Grey site rules) | 90/98 (92%); 0 wrong, 8 no answer |
| Polarity on names with no good direction (context, neutral, unclear) | 23 names | claims a direction on all | abstains on 17/23 (74%) |
| Polarity accuracy (per study, measure in the abstract) | 130 studies | 58/58 (100%) (implied, only where it can be derived) | 127/130 (98%) |
| Net direction, per study (exact, 5 classes) | 132 abstracts | 116/132 (88%) | 120/132 (91%) |
| Net direction, body of evidence (table verdict vs the review; exact) | 73 pairs | 44/73 (60%) | 64/73 (88%) |
| Same, claim level (benefit / harm / no claim) | 73 pairs | 53/73 (73%) | 69/73 (95%) |
| False-benefit rate on null reviews (review says no clear difference; table says improved) | 30 pairs | 7/30 (23%) | 0/30 (0%) |
| False benefit, null or uncertain reviews | 44 pairs | 13/44 (30%) | 0/44 (0%) |
| False benefit reading the null review's own abstract | 31 abstracts | 0/31 (0%) | 0/31 (0%) |
| Benefit recall, body of evidence | 25 pairs | 20/25 (80%) | 23/25 (92%) |
| Harm recall, per study (label worsens) | 10 abstracts | 9/10 (90%) | 9/10 (90%) |
| Harm precision, per study | 9/9 (100%) | 9/9 (100%) | |
| Harm recall, body of evidence | 2 pairs | 1/2 (50%) | 1/2 (50%) |
| The review itself among the studies read | 73 pairs | 22 | 64 |
Dr. Grey stored rows (old = origin/main outcomeHelpers.ts; new = interpret_row with the polarity resource; labels from the linked abstracts):
| Metric | Rows | Site today | New |
|---|---|---|---|
| Shown as "Worsens" when the evidence is not a harm | 42 | 28/42 (67%) | 2/42 (5%) |
| Verdicts shown (Improves/Worsens) that are wrong | 30/39 (77%) | 3/13 (23%) | |
| Real harms shown as "Worsens" (harm recall) | 4 | 2/4 (50%) | 4/4 (100%) |
| Net matches the label (rows with a readable label) | 38 | 9/38 (24%) | 11/38 (29%) (most rows hold no usable direction, so "unclear" is the honest display) |
| Polarity of the row's outcome name | 41 | 41/41 (100%) | 38/41 (93%) (3 no answer) |
The same rows answered by the engine (one table per row, supplement + the row's outcome name, no date limit):
| Metric | Rows | Old engine | New engine |
|---|---|---|---|
| Net matches the label (exact) | 38 | 17/38 (45%) | 14/38 (37%) |
| Claim level | 38 | 23/38 (61%) | 22/38 (58%) |
| False benefit (label no clear difference or mixed) | 14 | 4/14 (29%) | 2/14 (14%) |
| Benefit recall | 20 | 14/20 (70%) | 10/20 (50%) |
| Harm recall | 4 | 0/4 (0%) | 0/4 (0%) |
Contested labels left out: 6 study cases, 3 pairs, 1 rows, 5 names.
Reading: the held-out result is mixed, and the sets are small (18 null pairs, 6 harm abstracts).
- Fixed: false benefit against null reviews, harm recall per study, the review being among the studies read, and Dr. Grey rows shown as "Worsens" from a bare stored direction.
- Worse on held-out: exact body-level direction (7 of the 11 mismatches are "uncertain" where the label says "no clear difference"; one rule choice, switched twice on dev, decides this), polarity of outcome names by the rules (fewer wrong answers, many more with no answer; plain Qwen zero-shot scores 74/78 there), and benefit recall on Dr. Grey rows answered by the engine (the engine follows a review over the row's cited trials; which is right was not tested).
- Tuning disclosure: four body-level and three study-level configurations were scored on dev before freezing; several rules rest on fewer than ten dev cases; dev was enriched with known failures. Held-out was scored once.
- Cost per table at list price: $0.0066 before, $0.0107 after; p50 6.7 s before, 11.2 s after. The full write-up, method and failure types are in our internal notes.