Skip to content
decosa

The full write-up behind the numbers on the tool’s page, as the team that built it wrote it: data, method, results and caveats. Internal names are removed; nothing else is edited.

Eval: See what studies found for a supplement (use case 183, what-studies-found)

29 Sep 2026. Engine decosa_api/studies, rubric studies-rubric-2, Qwen3.8-27B. The run used our server's direct route, which serves the same weights as the gateway. Cost is at list price ($0.30/M input, $1.50/M output) from token counts. Wall time is measured on a pre-release server under load (3 tables at once), live PubMed included.

What is measured

  1. Verdict against Cochrane. For each supplement-outcome pair where a Cochrane review states a conclusion, we check whether the table's body direction and band match the review's direction (improved / no difference / worsened / uncertain) and its GRADE certainty.
    • Each table only sees studies published up to the review's year.
    • Two modes: A, as deployed (trials plus their meta-analyses and reviews, so the Cochrane review itself can be among the rows), and B, primary trials only.
  2. Every row against its own source. The grounding block (use case 17) judges every row's finding against that study's abstract automatically. Code checks every quote word for word. A blind judge checks a sample of rows field by field.
  3. Cost and time per table.

Labels came from blind sub-agents (Claude Code Opus 5.5) that read only the Cochrane abstracts and never saw engine output: evaldata/what-studies-found/set1.json and set2-heldout.json. They hold the review PMID, direction and certainty; abstract text is left out. There are two sets:

  • Set 1: 40 pairs.
  • Set 2: 31 held-out pairs from different reviews, made after rubric 2 was written.

Runner: evaldata/what-studies-found/eval_run.py, then summarize.py.

Tuning disclosure

Rubric 1 overstated certainty on set 1: 11 of the 24 pairs Cochrane rated low or very low came out "high". After seeing that, we wrote rubric 2:

  • the newest agreeing review's stated GRADE certainty caps the band;
  • without one, the band is at most "moderate", because abstracts can't show risk of bias;
  • a "very low" band states no direction.

Rubric 2 also extracts stated_certainty and drops joining words from PubMed terms. One set-1 pair returned 0 hits for "respiratory infections in young children", and 101 without "in".

Set 1 is therefore seen data for rubric 2. Set 2 is the honest number.

Results

Set Mode Pairs Direction agrees Band exact Band within 1 Rubric 1 on the same rows (dir / exact / within 1)
Set 2 (held-out) A 31 16 (52%) 14 (45%) 26 (84%) 16 / 7 / 18
Set 2 (held-out) B 31 18 (58%) 10 (32%) 24 (77%) 16 / 7 / 22
Set 1 (seen) A 40 26 (65%) 20 (50%) 33 (83%) 23 / 10 / 24
Set 1 (seen) B 40 15 (38%) 16 (40%) 31 (78%) 14 / 15 / 30
  • Where direction disagrees (set 2, A): 4 no_difference→improved, 3 uncertain→improved, 2 improved→no_difference and 2 uncertain→no_difference. The typical miss is a review that concludes "little or no difference" or "uncertain" while the abstracts the search returns read positive: small positive trials and older, optimistic meta-analyses.
  • Opposite directions (improved against no_difference or worsened): 6 of 31 (A) and 1 of 31 (B).
  • Rubric 2 against rubric 1: band agreement doubles on the held-out set (exact 7→14, within one 18→26). Direction is unchanged.

Rows against their sources

  • Grounding block (every row): set 2 A, 118 supported and 10 partial of 128 rows; set 1 A, 188 supported and 10 partial of 198. Rows judged unsupported or contradicted are dropped before grading and listed as excluded.
  • Quotes found word for word by code: 116 of 128 (set 2 A). A quote that fails the check is removed, so no row shows an unverified quote. Quotes are at most 15 words.
  • Blind judge (Claude Code Opus 5.5 sub-agent). It checked 100 rows sampled from set 1, stratified by direction, and saw only the abstract and the row; it never saw labels or grounding verdicts.
Field Result (of 100) Main errors
Relevant to the question 92 Dietary-intake or observational meta-analyses; mixed products; the supplement in both arms
Design 88 correct, 8 unclear, 4 wrong Mostly "meta-analysis" for a review that didn't pool
n 82 correct, 18 not stated, 0 wrong
Direction 96 correct
Finding 93 supported, 7 partly, 0 unsupported or contradicted
Quote 99 verbatim, 1 empty
  • Grounding against the judge: the grounding block flagged 3 of the judge's 7 "partly" rows as partial. It also flagged 6 rows as partial that the judge found fully supported, so it is conservative.

Cost and time per table (8 studies, 17–20 model calls)

  • Set 2 A: p50 $0.0068 and 8.7 s; p95 $0.0078 and 11.4 s.
  • Set 1 A: p50 $0.0071 and 9.1 s; p95 $0.0084 and 12.0 s.

That is list price, direct route. Gateway runs are metered at the same price.

What this means

  • The table is accurate enough to publish with review: rows cite real PMIDs, directions are right 96% of the time, findings hold up against the abstract, and quotes are verbatim.
  • The one-line verdict is not a substitute for a systematic review. On held-out pairs it agrees with Cochrane's direction about half the time, and it tends to read small positive trials as a benefit.
  • Products should lead with the table, show the band as "our rubric over the abstracts found", and link the rubric. Dr. Grey's proposals require human review before apply.

Checkable properties of the sample run (rehearsal/what-studies-found)

  1. Probiotics and antibiotic-associated diarrhea: the direction is improved, the label Improves, decided by a systematic review.
  2. The band is high, moderate or low (the certainty the best review states).
  3. At least 4 studies are counted, and the table carries a harm flag (none, possible or reported).
  4. Every row links to https://pubmed.ncbi.nlm.nih.gov/<PMID>/.
  5. The signed record verifies at /record/verify.
  6. Melatonin and sleep onset latency: the label is Mixed results, the claim none, the reasons say the reviews disagree, and a row's population is shift workers (the Cochrane review that takes best-review standing).
  7. A dosing question is refused with 400 ("no dosing").

What would move the numbers

  • Full text from PMC Open Access (CC-BY), so results reported only in the paper are read.
  • Deduplicating trials that a meta-analysis in the same table already pools.
  • A review-first mode that leads with the newest systematic review's conclusion.
  • Our own extractor (M-ev, page 82) trained on PMC-OA abstracts.

Harm read asymmetrically, wider review search, fresh split C (30 Sep 2026, the pre-release branch, frozen at b9108c64)

Prompts studies-extract-3 + studies-harm-1, net rule studies-net-2, rubric studies-rubric-4, against the engine that was live (6cabc7de: prompt 1, rubric 2). This supersedes the "Direction and meaning" section below: that version fixed the false benefits but lost harms (per study 26/27 -> 22/27 on its held-out split), so it was not merged. What changed since:

  • Harm is its own pass. Every abstract gets a second model call that looks only for harm: a worse result for the outcome (reported / possible / none, with the sentence, checked word for word in the abstract) and what the abstract says about adverse events in general. A reported harm makes the row "worsens" whatever the exposure type (a trial, a review, an association with intake or blood levels); it never overrules a statistic of the main reading, it flags the row instead. The table carries harm_signal ("Possible harm: check the studies") when any study read reports a worse result.
  • Add-on comparators count. "Chemotherapy alone", "vitamin D alone": both groups got the same treatment and one also got the supplement. Only a head-to-head comparison with a different treatment is set aside.
  • Review search. Supplement name variants and synonyms, the outcome with and without its condition, its other names from the polarity resource, reviews with the supplement in the title, the newest Cochrane reviews; up to 5 reviews always read (3 Cochrane, the newest version of each), added on top of the studies asked for.

Fresh held-out split C (never used for tuning; scored once)

Body of evidence (one table per review; the table's verdict against the review's), n = 289

Metric Old engine (live) New engine
Verdict, three-way (benefit / harm / no claim): MERGE RULE 189/289 (65%, 95% CI 60 to 71%) 231/289 (80%, 95% CI 75 to 84%)
False benefit on null reviews (label: no clear difference): MERGE RULE 8/61 (13%, 95% CI 7 to 24%) 1/61 (2%, 95% CI 0 to 9%)
Harm surfaced (verdict Worsens, or the harm flag): MERGE RULE 18/40 (45%, 95% CI 31 to 60%) 36/40 (90%, 95% CI 77 to 96%)
Harm recall, strict (verdict Worsens only) 18/40 (45%, 95% CI 31 to 60%) 28/40 (70%, 95% CI 55 to 82%)
False harm verdict (label is not a harm, verdict Worsens): guard 10/249 (4%, 95% CI 2 to 7%) 6/249 (2%, 95% CI 1 to 5%)
Harm verdict or flag on a label that is not a harm (the price of the flag) 10/249 (4%, 95% CI 2 to 7%) 46/249 (18%, 95% CI 14 to 24%)
False benefit, null or too-weak-to-tell reviews 36/140 (26%, 95% CI 19 to 34%) 11/140 (8%, 95% CI 4 to 14%)
Benefit recall 56/82 (68%, 95% CI 58 to 77%) 56/82 (68%, 95% CI 58 to 77%)
Verdict, five-way exact (both labellers agree on the five-way label) 135/282 (48%, 95% CI 42 to 54%) 180/282 (64%, 95% CI 58 to 69%)
Review found: the labelled review among the studies read 161/290 (56%, 95% CI 50 to 61%) 256/290 (88%, 95% CI 84 to 91%)
the labelled review among the counted rows 129/290 (44%, 95% CI 39 to 50%) 205/290 (71%, 95% CI 65 to 76%)
Cochrane pairs: the labelled Cochrane review among the studies read 87/154 (56%, 95% CI 49 to 64%) 143/154 (93%, 95% CI 88 to 96%)
Cochrane pairs: among the counted rows 64/154 (42%, 95% CI 34 to 49%) 113/154 (73%, 95% CI 66 to 80%)
"Too few studies to tell" 30/290 (10%, 95% CI 7 to 14%) 22/290 (8%, 95% CI 5 to 11%)

Per study (the engine reads the review's own abstract), n = 331

Metric Old engine (live) New engine
Verdict, three-way (benefit / harm / no claim) 280/331 (85%, 95% CI 80 to 88%) 290/331 (88%, 95% CI 84 to 91%)
False benefit on null reviews (label: no clear difference) 1/70 (1%, 95% CI 0 to 8%) 0/70 (0%, 95% CI 0 to 5%)
Harm surfaced (verdict Worsens, or the harm flag) 44/62 (71%, 95% CI 59 to 81%) 60/62 (97%, 95% CI 89 to 99%)
Harm recall, strict (verdict Worsens only) 44/62 (71%, 95% CI 59 to 81%) 51/62 (82%, 95% CI 71 to 90%)
False harm verdict (label is not a harm, verdict Worsens): guard 3/269 (1%, 95% CI 0 to 3%) 8/269 (3%, 95% CI 2 to 6%)
Harm verdict or flag on a label that is not a harm (the price of the flag) 3/269 (1%, 95% CI 0 to 3%) 38/269 (14%, 95% CI 10 to 19%)
False benefit, null or too-weak-to-tell reviews 12/153 (8%, 95% CI 5 to 13%) 1/153 (1%, 95% CI 0 to 4%)
Benefit recall 69/84 (82%, 95% CI 73 to 89%) 68/84 (81%, 95% CI 71 to 88%)
Verdict, five-way exact (both labellers agree on the five-way label) 229/324 (71%, 95% CI 66 to 75%) 258/324 (80%, 95% CI 75 to 84%)

Set: 389 abstracts labelled by both blind labellers; 331 are supplement reviews with a result for the outcome by both (8 more left out because the labellers disagree on benefit / harm / no claim); 290 body pairs. Labels of the body pairs (labeller A): 83 improves, 79 uncertain, 61 no clear difference, 40 worsens, 27 mixed. By source: 151 cochrane-pool, 75 harm-topic, 36 safety-pool, 28 general-pool. The two labellers give the same five-way label in 324/339 (96%, 95% CI 93 to 97%) of the cases both call relevant.

One more pair could not be run for the new engine (PubMed answered HTTP 429 while another of our runs shared the address); it is left out of both columns. Its label is "improves" and the old engine's verdict on it was "unclear", so counting it against the new engine changes no comparison.

Merge rule (set by the lead before the run): three-way verdict, false benefit on null reviews and harm recall each no worse than the old engine, at least two better. Three-way 189/289 -> 231/289; false benefit on null reviews 8/61 -> 1/61; harm surfaced 18/40 -> 36/40 (strict, the verdict itself: 18/40 -> 28/40). Met.

Where the 40 labelled harms went (new engine): 28 have the verdict "Worsens", 8 carry the "possible harm" flag under another verdict, 4 are not surfaced (3 Too few studies to tell; 1 No clear difference; in 3 of them the labelled review was not among the studies read).

The price of the flag: 46/249 tables whose labelled review reports no harm carry a harm verdict or flag (27 harm-topic, 11 cochrane-pool, 6 safety-pool, 2 general-pool: most are harm topics, where other studies read do report a worse result). A false harm VERDICT: 10/249 old, 6/249 new.

Remaining failure types (new engine, body, label -> verdict): improves -> mixed: 13; uncertain -> improves: 10; improves -> unclear: 8; worsens -> unclear: 5; worsens -> no_clear_difference: 5; mixed -> improves: 4; mixed -> worsens: 4; improves -> no_clear_difference: 4; worsens -> mixed: 2; improves -> worsens: 1; no_clear_difference -> improves: 1; no_clear_difference -> worsens: 1. By where the labelled review went: the labelled review counted, another study decided or it was read differently: 25; the labelled review was not read: 13; read, an association with intake or levels (no harm): 9; read, called not about this supplement and outcome: 4; read, set aside as a head-to-head comparison: 3; read, no usable result: 2; read, its finding did not hold up against its abstract: 2. Benefits missed (26 of 82, the same count as the old engine): Mixed results: 13; Too few studies to tell: 5; No clear difference: 4; Unclear: evidence too weak to tell: 3; Worsens: 1. Per study, a false harm verdict rose from 3/269 to 8/269: label mixed, decided by harm reading, exposure given: 1; label mixed, decided by harm reading, exposure intake_or_levels: 3; label mixed, decided by polarity, exposure intake_or_levels: 2; label uncertain, decided by polarity, exposure given: 1; label improves, decided by polarity, exposure intake_or_levels: 1.

Review-found rate: the labelled review is among the studies read in 256/290 tables (old 161/290), and among the counted rows in 205/290 (old 129/290). For pairs whose labelled review is a Cochrane review: read 143/154 (old 87/154), counted 113/154 (old 64/154). In 222 tables a Cochrane candidate was found and its newest version was read; in 67 no Cochrane review matched the search; in 0 one was found and not read.

Method. 389 review abstracts, none used in any earlier split: every unused Cochrane supplement review from 2005 on with a supplement word in the title, the top 2 PubMed reviews for harm topics written down before any search (some are expected nulls), a seeded sample of safety-titled supplement meta-analyses and a seeded sample of general supplement meta-analyses of randomised trials. The supplement and outcome of the unframed cases were set by the model (no verdict asked) before labelling. Two blind Claude Code Opus labellers, in separate folders, each given the same supplement and outcome and neutral case ids; the second never saw the first's fields. A case where they disagree on benefit / harm / no claim is left out. The engine was frozen (b9108c64) and the scoring rule written down before any engine output on this split existed. Both engines ran in-process on our server's direct route (same weights as the gateway), temperature 0, one shared PubMed cache, tables limited to studies published up to the review's year.

Cost and time per table (8 studies asked for, up to 11 read): new 21.3 model calls and $0.0142 at list price (old 10.8 calls, $0.0056); wall time p50 30 s, p95 56 s with 4 tables at a time and another run sharing the model (old p50 15 s).

Dev (the earlier fresh split, read before this work, so tuning data): kept with the eval harness, not reported here as a result.

The samples (30 Sep 2026): chosen for a stable verdict, engine unchanged

The engine merged as frozen. The flagship sample changed: melatonin and sleep onset latency reads "Mixed results" on this engine (a Cochrane review about shift workers takes best-review standing and the general reviews disagree), so the first sample is now a pair from the held-out split that the engine got right, whose verdict was the same in every run below and matches its Cochrane review. Melatonin stays as the second sample and shows the disagreement as it is. Runs had no date limit, as on the site.

Sample Supplement and outcome In process, 4 rounds Through the API, 4 runs Deciding review
probiotics-antibiotic-diarrhea probiotics / antibiotic-associated diarrhea Improves / moderate: 4 of 4 Improves / moderate: 4 of 4 PMID 31039287 (Cochrane)
melatonin-sleep-onset melatonin / sleep onset latency Mixed results / very low: 4 of 4 Mixed results / very low: 4 of 4 PMID 25113164 (Cochrane)
beta-carotene-lung-cancer beta-carotene / lung cancer Worsens / moderate: 4 of 4 Worsens / moderate: 4 of 4 PMID 37702300 (Cochrane)
caffeine-blood-pressure caffeine / blood pressure Mixed results / very low: 4 of 4 Mixed results / very low: 4 of 4 PMID 38057002
vitamin-d-asthma vitamin D / severe asthma exacerbations No clear difference / high: 4 of 4 No clear difference / high: 4 of 4 PMID 36744416 (Cochrane)
omega3-depression omega-3 / depressive symptoms Unclear: evidence too weak to tell / very low: 4 of 4 Unclear: evidence too weak to tell / very low: 4 of 4 PMID 39564892 (Cochrane)
cranberry-uti cranberry / urinary tract infection Improves / moderate: 4 of 4 Improves / moderate: 4 of 4 PMID 37068952 (Cochrane)

Candidates run and not used (in process, 4 rounds each): iron / anaemia (No clear difference / low: 4 of 4); creatine / muscle hypertrophy (Improves / low: 3 of 4; Mixed results / very low: 1 of 4); St John's wort / depressive symptoms (Improves / moderate: 4 of 4); psyllium / fasting blood sugar (Improves / moderate: 4 of 4); curcumin / fasting blood glucose (Improves / moderate: 4 of 4); iron / restless legs syndrome severity (Improves / moderate: 4 of 4); myo-inositol / gestational diabetes (Improves / low: 4 of 4); soluble fiber / serum lipid profile (Improves / moderate: 4 of 4); zinc / common cold duration (Improves / low: 4 of 4); vitamin C / common cold duration (Improves / moderate: 4 of 4); magnesium / diarrhoea (Worsens / moderate: 4 of 4); vitamin A / hip fracture (Worsens / moderate: 4 of 4); berberine / gastrointestinal adverse events (Mixed results / very low: 4 of 4); omega-3 fatty acids / cognitive function (Mixed results / very low: 4 of 4); vitamin A / all-cause mortality (No clear difference / high: 4 of 4).

The 28 API runs cost $0.0182 to $0.0231 each at list price and took 11.8 to 18.2 s (direct route, one table at a time).

The machine-readable scores for split C (the scorer's own output) are in docs/evals/what-studies-found-split-c.json; the site's figures are generated from it.

Direction and meaning (30 Sep 2026, on a pre-release build, measured at 4f2101b3; superseded by the section above, never merged on its own)

Prompt studies-extract-2, net rule studies-net-1, rubric studies-rubric-3, polarity resource decosa-polarity-1, against the engine above (prompt 1, rubric 2). Labels by a blind Claude Code Opus sub-agent from the source text, a second blind labeller on every disagreement (cases where the two differ are left out). Old and new ran on the direct route with one shared PubMed cache. Tables below are generated from the score files.

Held-out split (never used for tuning)

Metric Set Old New
Polarity accuracy (outcome names; labelled higher/lower) 78 names 61/78 (78%) (Dr. Grey site rules) 56/78 (72%); 4 wrong, 18 no answer
Polarity on names with no good direction (context, neutral, unclear) 15 names claims a direction on all abstains on 10/15 (67%)
Polarity accuracy (per study, measure in the abstract) 67 studies 31/31 (100%) (implied, only where it can be derived) 66/67 (99%)
Net direction, per study (exact, 5 classes) 68 abstracts 57/68 (84%) 60/68 (88%)
Net direction, body of evidence (table verdict vs the review; exact) 40 pairs 33/40 (82%) 29/40 (72%)
Same, claim level (benefit / harm / no claim) 40 pairs 35/40 (88%) 37/40 (92%)
False-benefit rate on null reviews (review says no clear difference; table says improved) 18 pairs 2/18 (11%) 0/18 (0%)
False benefit, null or uncertain reviews 29 pairs 3/29 (10%) 1/29 (3%)
False benefit reading the null review's own abstract 18 abstracts 0/18 (0%) 0/18 (0%)
Benefit recall, body of evidence 10 pairs 9/10 (90%) 9/10 (90%)
Harm recall, per study (label worsens) 6 abstracts 5/6 (83%) 6/6 (100%)
Harm precision, per study 5/5 (100%) 6/6 (100%)
Harm recall, body of evidence 1 pairs 0/1 (0%) 0/1 (0%)
The review itself among the studies read 40 pairs 13 32

Dr. Grey stored rows (old = origin/main outcomeHelpers.ts; new = interpret_row with the polarity resource; labels from the linked abstracts):

Metric Rows Site today New
Shown as "Worsens" when the evidence is not a harm 42 24/42 (57%) 2/42 (5%)
Verdicts shown (Improves/Worsens) that are wrong 27/35 (77%) 4/14 (29%)
Real harms shown as "Worsens" (harm recall) 4 2/4 (50%) 4/4 (100%)
Net matches the label (rows with a readable label) 37 8/37 (22%) 13/37 (35%) (most rows hold no usable direction, so "unclear" is the honest display)
Polarity of the row's outcome name 42 42/42 (100%) 40/42 (95%) (2 no answer)

The same rows answered by the engine (one table per row, supplement + the row's outcome name, no date limit):

Metric Rows Old engine New engine
Net matches the label (exact) 37 20/37 (54%) 13/37 (35%)
Claim level 37 24/37 (65%) 17/37 (46%)
False benefit (label no clear difference or mixed) 12 4/12 (33%) 3/12 (25%)
Benefit recall 21 15/21 (71%) 9/21 (43%)
Harm recall 4 1/4 (25%) 0/4 (0%)

Contested labels left out: 2 study cases, 0 pairs, 5 rows, 1 names.

Dev split (rules and prompt were tuned on these)

Metric Set Old New
Polarity accuracy (outcome names; labelled higher/lower) 98 names 78/98 (80%) (Dr. Grey site rules) 90/98 (92%); 0 wrong, 8 no answer
Polarity on names with no good direction (context, neutral, unclear) 23 names claims a direction on all abstains on 17/23 (74%)
Polarity accuracy (per study, measure in the abstract) 130 studies 58/58 (100%) (implied, only where it can be derived) 127/130 (98%)
Net direction, per study (exact, 5 classes) 132 abstracts 116/132 (88%) 120/132 (91%)
Net direction, body of evidence (table verdict vs the review; exact) 73 pairs 44/73 (60%) 64/73 (88%)
Same, claim level (benefit / harm / no claim) 73 pairs 53/73 (73%) 69/73 (95%)
False-benefit rate on null reviews (review says no clear difference; table says improved) 30 pairs 7/30 (23%) 0/30 (0%)
False benefit, null or uncertain reviews 44 pairs 13/44 (30%) 0/44 (0%)
False benefit reading the null review's own abstract 31 abstracts 0/31 (0%) 0/31 (0%)
Benefit recall, body of evidence 25 pairs 20/25 (80%) 23/25 (92%)
Harm recall, per study (label worsens) 10 abstracts 9/10 (90%) 9/10 (90%)
Harm precision, per study 9/9 (100%) 9/9 (100%)
Harm recall, body of evidence 2 pairs 1/2 (50%) 1/2 (50%)
The review itself among the studies read 73 pairs 22 64

Dr. Grey stored rows (old = origin/main outcomeHelpers.ts; new = interpret_row with the polarity resource; labels from the linked abstracts):

Metric Rows Site today New
Shown as "Worsens" when the evidence is not a harm 42 28/42 (67%) 2/42 (5%)
Verdicts shown (Improves/Worsens) that are wrong 30/39 (77%) 3/13 (23%)
Real harms shown as "Worsens" (harm recall) 4 2/4 (50%) 4/4 (100%)
Net matches the label (rows with a readable label) 38 9/38 (24%) 11/38 (29%) (most rows hold no usable direction, so "unclear" is the honest display)
Polarity of the row's outcome name 41 41/41 (100%) 38/41 (93%) (3 no answer)

The same rows answered by the engine (one table per row, supplement + the row's outcome name, no date limit):

Metric Rows Old engine New engine
Net matches the label (exact) 38 17/38 (45%) 14/38 (37%)
Claim level 38 23/38 (61%) 22/38 (58%)
False benefit (label no clear difference or mixed) 14 4/14 (29%) 2/14 (14%)
Benefit recall 20 14/20 (70%) 10/20 (50%)
Harm recall 4 0/4 (0%) 0/4 (0%)

Contested labels left out: 6 study cases, 3 pairs, 1 rows, 5 names.

Reading: the held-out result is mixed, and the sets are small (18 null pairs, 6 harm abstracts).

  • Fixed: false benefit against null reviews, harm recall per study, the review being among the studies read, and Dr. Grey rows shown as "Worsens" from a bare stored direction.
  • Worse on held-out: exact body-level direction (7 of the 11 mismatches are "uncertain" where the label says "no clear difference"; one rule choice, switched twice on dev, decides this), polarity of outcome names by the rules (fewer wrong answers, many more with no answer; plain Qwen zero-shot scores 74/78 there), and benefit recall on Dr. Grey rows answered by the engine (the engine follows a review over the row's cited trials; which is right was not tested).
  • Tuning disclosure: four body-level and three study-level configurations were scored on dev before freezing; several rules rest on fewer than ten dev cases; dev was enriched with known failures. Held-out was scored once.
  • Cost per table at list price: $0.0066 before, $0.0107 after; p50 6.7 s before, 11.2 s after. The full write-up, method and failure types are in our internal notes.