Skip to content
decosa

The full write-up behind the numbers on the tool’s page, as the team that built it wrote it: data, method, results and caveats. Internal names are removed; nothing else is edited.

Eval: EU trial lay summary with number grounding (66)

Run 26 Sep 2026 on our server, through the shared hosted gateway (Qwen3.8-27B NVFP4, temperature 0, thinking off), branch the pre-release branch. Runner: scripts/eval_laysummary.py. Results: docs/evals/eu-trial-lay-summary/<split>/.

Data

Twenty real trials with results posted on ClinicalTrials.gov (API v2 records fetched 26 Sep 2026, stored gzipped in trials/; cited by NCT number). All phase 3, randomised, with an EudraCT number, from 19 sponsors and many areas (insulin and type 2 diabetes, COVID-19, hand eczema, Crohn's disease, fibromyalgia, paediatric MS, a children's flu vaccine, thyroid eye disease, uterine fibroids, CMV after kidney transplant, fibrinogen in surgery, overactive bladder in men and in children, PNH, Dravet-type seizures, atopic dermatitis, psoriatic arthritis, influenza vaccine).

split trials used for
dev 3 (the samples: NCT04880850, NCT04362137, NCT05355818) writing the prompts and the checker
test 12 run once with everything frozen; its results are reported as the headline, then its misses were fixed
test2 5 kept fresh; run after the fixes the test split led to, to check them on data they were not made on

Before the split I looked at the trial list and at each record's group and analysis structure (to pick two-arm and multi-arm trials); I did not draft or check any test or test2 trial before its run.

Published lay summaries were not used: the ones in CTIS are PDFs behind the portal's document store, and sponsor sites publish them under their own terms. Comparing our drafts with real ones is not done (see Limits).

What was measured

  1. Drafts (/laysummary/draft, one per trial): the model's first-draft number accuracy (before the repair pass), what the checker left after the repair, grounding verdicts, Annex V coverage, readability, receipts and time.
  2. Planted errors (/laysummary/check, grounding off, so this measures the code checks): from each final draft, sentences the checker had passed were changed one at a time: a number changed (±1 or ±3 to 10%, or the last decimal), a percentage changed (±1 to 3 points or the last decimal), two groups' figures swapped, a date moved a year, a "could be chance" result restated as a difference, a comparison flipped, "only" put before a serious side-effect count, "no one in <group> had a serious side effect" when the table has some, and one group's serious side-effect sentence deleted. Caught = the changed sentence comes back fail or check for a reason of the right kind (a number status for number changes, direction for a flip, hedge_missing, minimising, contradicts_table, and for a deleted group an ae_incomplete document flag). Each draft also went through unchanged: every flag on it is a candidate false flag, adjudicated by me. Everything was run twice: with the draft's [C12] citations, and with them stripped (a sponsor's own draft).
  3. Wording errors through the grounding judge (/laysummary/check, grounding on): in 48 sentences of the test drafts with no number in play, one word swapped for a wrong one (tablet/injection, once a day/once a week, 18/65 years old, could/could not take part, placebo/standard treatment, children/adults). Only the changed sentence's section is sent. Caught = the judge says contradicted, unsupported or partial.
  4. Clean drafts through the grounding judge after the last fixes (test2, recheck).
  5. My own review of 10 test summaries, read as a medical writer doing QC would (below). This is one reader, the same agent that built the tool, not an independent reviewer or a lay-reader test.

Results

Drafts (test, 12 trials; test2, 5 trials)

test test2
Numbers in the model's first drafts 642 289
... wrong or in no source (before the repair) 5 (0.8%) 1 (0.3%)
... right but not in the cells cited, or uncited 12 19
Sections rewritten once 12 of 72 6 of 30
Numbers in the final drafts 655 301
... traced to a cell, fact, protocol sentence or arithmetic 646 284
... wrong 0 1 (a false flag: a dose inside a group name, fixed since)
Sentences 812 347
... flagged fail / check 12 / 36 7 / 27
... gaps for the sponsor ([Sponsor to add]) 68 27
Narrative sentences judged supported / partial / unsupported / contradicted / no claim 584 / 30 / 7 / 3 / 36 249 / 12 / 3 / 3 / 18
Flesch-Kincaid grade (whole summary): min / median / max 4.4 / 6.1 / 8.0 4.9 / 6.3 / 6.8
Summaries above grade 8 0 0
Model calls (receipts) per draft, mean 62.5 63.8
Wall time per draft, median (range), 3 in parallel on the shared gateway 115 s (88-289 s) 300 s (219-363 s)

Annex V coverage, every test and test2 draft: elements 5-8 and 10 covered in all 17; element 3 (general information) in 14, partial in 3 (the "why" was not stated in words the check looks for); elements 1, 2, 4 and 9 always "needs input": the registry does not hold the EU trial number, the sponsor's contact details, the participants per Member State, in the Union and in third countries, or follow-up plans. The draft leaves marked gaps for them; with the sponsor's inputs they are filled (tested in tests/test_laysummary.py, not in this eval).

Planted errors (code checks, grounding off)

dev (3), after all fixes test, frozen (12) test2, after the test fixes (5) test, after all fixes test2, after all fixes
Caught, with citations 130 / 132 555 / 577 (96.2%) 212 / 217 (97.7%) 576 / 577 211 / 217
number changed 86 / 88 365 / 368 132 / 136 367 / 368 131 / 136
percentage changed 29 / 29 135 / 137 53 / 54 137 / 137 53 / 54
two groups' figures swapped 1 / 1 7 / 8 none planted 8 / 8 none planted
date moved 3 / 3 15 / 15 8 / 8 15 / 15 8 / 8
"could be chance" restated as a difference 2 / 2 1 / 4 none planted 4 / 4 none planted
comparison flipped 2 / 2 4 / 9 4 / 4 9 / 9 4 / 4
"only" before a side-effect count 3 / 3 10 / 12 5 / 5 12 / 12 5 / 5
"no one ... had a serious side effect" when some did 2 / 2 6 / 12 5 / 5 12 / 12 5 / 5
one group's serious side effects deleted 2 / 2 12 / 12 5 / 5 12 / 12 5 / 5
Clean sentences flagged, with citations 0 / 186 10 / 812 16 / 347 15 / 812 6 / 347
... of which false flags (my adjudication) 0 9 / 812 (1.1%) 16 / 347 (4.6%) 6 / 812 1 / 347
Caught, citations stripped 69 / 132 315 / 577 (54.6%) 114 / 217 320 / 577 114 / 217
Clean sentences flagged, citations stripped 0 / 186 15 / 812 3 / 347 23 / 812 8 / 347

The frozen test run is the held-out number. Its 22 misses: group names with punctuation ("Double-Blind Treatment Period: Placebo") broke the comparison pattern; "no one in <group>" claims were only caught when short; a "not due to chance" sentence nearby counted as a hedge; a frequency over a large base ("29 out of 2255") was accepted as a nearby proportion; a group's size was accepted as any number in the sentence ("57 people finished" when 57 started); a figure swapped to the other group was not checked against the group named; and the eval's own "only" was put at the start of a sentence, where the check lets "only in the first period" pass. Those were fixed and each has a unit test; test2 was then run fresh on the fixed code. Test2's 16 flagged clean sentences came from one trial where the model shortened group names that carry a dose ("the ZX008 0.2 mg/kg/Day group"), so the dose was read as a claim; that was fixed after test2, so the last two columns are no longer held out.

After all fixes, the flags on clean drafts include two checks added after the eval (not in the frozen or test2 runs): "was not due to chance" (a statistical test can only say a difference is unlikely to be chance; 7 test and 5 test2 sentences) and a proportion written as a decimal ("0.442 of the group"; 2 test sentences). I count those as true flags: the recommendations ask for numbers of people or percentages, and the phrase overstates certainty. The false flags left are right numbers that sit in no cell the sentence cites ("1 tablet and 1 capsule twice a day", "within 52 weeks") and one "no people died in either group during the second period", where the check cannot tell periods apart.

Without citations the number check falls back to the tables of the sentence's own section, and about half the changed numbers happen to equal another cell there. A sponsor's own draft should cite, or be checked with the grounding judge on. Comparisons are only checked when the sentence cites the cells.

Wording errors through the grounding judge (test)

RESULT_CLAIMS

Clean drafts through the grounding judge, after the last fixes (test2)

RESULT_RECHECK

Flags on the 12 test drafts, adjudicated

The 48 sentences the drafts came back with as fail or check (code and grounding together), each read against the record: 9 point at a real problem (for example "Two groups took a dummy treatment" in a trial with one placebo group; 2.88 reported as the drop in the main outcome when it was urgency episodes; a median called an average; 289 "took letermovir" when 289 were analysed and 292 treated; the registry giving 22 countries in one place and 21 in another; a neutrophil criterion turned into white blood cells; a wrong plain-language gloss of "diffuse axonal injury"); 17 are debatable (a simplified criterion, "most common" read from the top rows, a route of administration the sources do not state); 22 are false flags (6 are the judge refusing plain definitions of placebo or a dummy treatment, 8 are right numbers not in the cited cells, 3 are the judge missing a zero-death table row or a unit). 22 false flags in 812 sentences is 2.7%. Two of the fixes made after the eval target these: a plain-language glossary source for the judge (placebo, randomised, double-blind, infusion, median, following the EU recommendations' Annex 1 wording), and the table's measure type (median, mean) in the judge's sources.

My review of 10 summaries

I read ten test summaries in full (NCT01277289, NCT02187471, NCT02201108, NCT03165617, NCT03298867, NCT03400943, NCT03443869, NCT03444324, NCT04409262, NCT04641975) as QC would, against the posted records. This is one reader who also built the tool; treat it as a sanity check, not a study.

  • Numbers: I found no wrong number in the final drafts. I read the trace of every side-effect and results number in four of them (97 numbers): each pointed to the right cell, and percentages were rounded correctly.
  • Side effects: every summary gave serious side effects and deaths for every group, with numbers, and said the tables count problems whatever their cause. None softened them. Two used "0 out of 110 (0%)" style lines, which are right.
  • Could be chance: every summary whose main analysis could be chance said so. The problem runs the other way: 4 of 10 wrote "the difference was not due to chance" for significant results, which overstates what a test shows. The checker did not flag this during the eval; a check for it was added afterwards.
  • What the checker missed (a writer would still fix these): "went down by 3.84 times" for 3.84 fewer urinations a day (a reader could take it as a factor); a secondary result where the placebo group improved more, with no word on which group did better; "the age spread was 14.9 years" for a standard deviation; a proportion given as "0.442 of the group"; a non-inferiority trial summarised as "could be due to chance" without saying the new medicine was shown to be not worse; generic closing sentences carrying stray "[Sponsor to add: ...]" brackets in two summaries.
  • Readability: grades 4.4 to 8.0. They read plainly, but multi-arm trials with long group names ("Placebo+Vilaprisan (B1)", "Double-Blind Treatment Period: Teriflunomide") are hard going; a writer would rename the groups.
  • Honesty of the gaps: every summary marked what the registry does not hold instead of inventing it.

Limits

  • The drafts were written by our model; the eval does not include summaries written by people, or published lay summaries. Comparing with CTIS lay summaries is the next step.
  • The planted errors are ours, generated in code from the drafts, one per sentence. Real errors can be subtler (the 3.84 "times" case) and several at once.
  • One judge of false flags and of the 10 summaries: the agent that built the checker. No lay-reader test, which the EU recommendations advise.
  • English drafting only. Numbers and readability are checked in five more languages (unit-tested), not evaluated.
  • ClinicalTrials.gov records only; CTIS and EudraCT tables were not read.
  • Times were measured under load from other workloads on the same gateway; they vary a lot.
  • Not "compliant": coverage says which of Annex V's parts are present, not whether they are adequate.

Expected properties of the sample run (for the rehearsal kit)

On the check bundle (rehearsal/eu-trial-lay-summary/, the icodec trial and a three-sentence side-effects section):

  1. the sources come back without a model call, and C114 is the icodec serious side-effect count (22 of 291);
  2. the correct glargine sentence is ok and the grounding judge supports it;
  3. the icodec sentence with 23 is a number mismatch whose nearest cell is C114;
  4. "No one in the trial died" raises a contradicts_table document flag;
  5. the signed record verifies, and fails once its draft hash is changed;
  6. every model call has a signed receipt.

Verdict

VERDICT