66 · Healthcare · Science and research · live
EU trial lay summary with number grounding
Eval results
Scored on a held-out or test splitRun 27 Sep 2026Eval write-up (decosa-api, access required)
- Planted errors caught, drafts with citations (code checks)555 / 577 (96.2%)test splitn = 57712 held-out ClinicalTrials.gov trials; numbers, percentages, dates, swapped groups, flipped comparisons, dropped hedges, softened or denied side effects.
- Correct sentences flagged (false flags), code checks9 / 812 (1.1%)test splitn = 812Adjudicated by the author.
- Planted errors caught after fixes, fresh trials212 / 217 (97.7%)held outn = 2175 trials kept aside while the test split's misses were fixed; 16 of 347 correct sentences flagged there, from one trial's shortened group names (fixed afterwards).
- Planted errors caught, citations stripped315 / 577 (54.6%)test splitn = 577
- Wrong numbers in the model's first drafts5 / 642 (0.8%)test splitn = 642None left after the check and one repair pass (0 / 655).
- Planted wording errors caught by the grounding judge29 / 45 (64%)test split
- Flesch-Kincaid grade, median (range)6.1 (4.4-8.0)test splitn = 12
- Planted errors caught, dev130 / 132dev (tuned on)n = 132The 3 sample trials, used while writing the prompts and the checker.
- Numbers in Member-State versions traced to the results, 6 languages (frozen)4,002 / 4,205 (95.2%)test splitn = 4,205The 12 test trials' English drafts translated into de, fr, es, it, nl, pl. Most untraced numbers were the checker misreading other languages (fixed afterwards: 4,014 / 4,077, no longer held out).
- Planted number errors in the translations caught2,426 / 2,432 (99.8%)test splitn = 2,432One digit changed in a sentence both checks had passed; 917 / 920 on the 5 test2 trials. Real errors the check found in the model's translations: 6 Polish sentences that dropped a '0 out of' count (test2) and one Spanish '30,000' left in English format.
Dataset
20 real phase 3 trials with results on ClinicalTrials.gov (19 sponsors, fetched 26 Sep 2026): dev 3, test 12 (run once, frozen), test2 5 (kept fresh for the fixes the test split led to). Drafts by our model; errors planted in code, one per sentence. Member-State versions: the test and test2 drafts translated into six languages on 27 Sep 2026.
Caveats
- The drafts and the planted errors are ours; no summaries written by people and no published lay summaries were checked.
- False flags and the 10-summary review were judged by the agent that built the checker; no independent reviewer and no lay-reader test.
- After the test run its misses were fixed; the test numbers after the fixes (576 / 577) are no longer held out, and the fresh test2 run (212 / 217) checks the fixes.
- ClinicalTrials.gov records only; CTIS and EudraCT tables were not read.
- Coverage says which Annex V parts are present, not whether they are adequate; it never says a summary complies.
- Member-State versions are machine translations with numbers checked in code; the optional meaning check caught all 6 known Polish "0 out of" drops and real errors such as "Vehicle Cream" as a cream for cars, but also raises false alarms (9 of 14 errors on a 240-sentence sample). Grammar is not checked; no native speaker has read them.
- The fifteen EU languages added on 27 Sep (Qwen3.8-27B route) were run on the 5 test2 trials in five of them (fi, el, hu, ro, ga): 1,604 of 1,624 numbers traced to the results, 1 wrong (Irish); real errors found were a dropped Finnish count and two wrong Irish month names. No native speaker has read any version.
Nightly smoke check
Loading the nightly status…
- Result
- pass
- Run
- 26 Sep 2026
- Latency, this run
- n/a
- p50 over passed runs
- 8.2 s
- Receipts
- 3
- Model calls
- n/a
- Tokens
- n/a
- Cost per run
- $0.001
Self-host verification
Verified on 26 Sep 2026: fresh clone, compose up, sample against local model servers
A fresh clone of a decosa-api pre-release build (not yet merged), the api image built from it with DECOSA_LAYSUMMARY_FETCH=0, run against the already-running local Qwen3.8-27B vLLM on the direct route. The rehearsal bundle passed 9 of 9 checks in 0.8 s, the smoke module passed in 0.9 s with 3 attested receipts, a full draft of the ruxolitinib sample took 18.4 s (59 attested receipts, 49 of 49 numbers traced, grade 5.7) and its record verified, and an NCT number was refused with fetching off. Torn down after. Model-server startup itself not re-verified.
Rehearsal bundle: eu-trial-lay-summary.zip (17 KB, 9 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.
Known limits
- Drafts in English; Member-State versions are machine translations with their numbers checked in code, not their wording.
- Reads ClinicalTrials.gov records; CTIS and EudraCT results tables are not read.
- Without citations in a draft, about half of the planted wrong numbers were missed (315 of 577 caught): cite the cells or turn the grounding judge on.
- The grounding judge can be wrong and is not fully repeatable on the shared gateway; one correct sentence was once called contradicted.
- Coverage says which Annex V parts are present, never that a summary meets Annex V.
Models and licences
Standard tier. Licence posture: permissive (Apache, MIT or BSD).
- Results normaliser, number tracing, comparison and hedge checks, side-effect completeness, word lists, Annex V coverage, readability, review file and signed record (no model; CPU)decosa-api lay summary (decosa_api/verticals/laysummary), importing numeric grounding (53) and table cells (57), the grounding judge (17) and the signed record (07)AGPL-3.0-or-later
- Drafts the six narrative sections with citations, rewrites a failed section once, and judges each sentence against the sourcesQwen3.8-27B (NVFP4)Apache-2.0
- Member-State versions: each sentence translated, its numbers checked against the English sentence and traced to the results cells in that language: German, French, Spanish, Italian, Dutch, Polish, Portuguese, CzechHy-MT2-7B (the language-pack block)Apache-2.0
- Member-State versions: the fifteen other EU languages (Swedish, Danish, Finnish, Greek, Romanian, Hungarian, Bulgarian, Croatian, Slovak, Slovenian, Lithuanian, Latvian, Estonian; Irish and Maltese as drafts)Qwen3.8-27B (the language-pack block's route for these languages)Apache-2.0
All quality evidence
Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.
Lite · check a draft in code, no GPU (3)
- Planted errors caught in drafts with citations (12 held-out trials): 555 / 577 (96.2%)docs/evals/eu-trial-lay-summary.md, test split, 26 Sep 2026
- Correct sentences flagged (false flags, same drafts): 9 / 812 (1.1%)docs/evals/eu-trial-lay-summary.md, test split, adjudicated by the author
- Planted errors caught with the citations stripped: 315 / 577 (54.6%)docs/evals/eu-trial-lay-summary.md, test split
Standard · one GPU for the model (hosted demo) (7)
- Wrong numbers in the model's first drafts (12 held-out trials): 5 / 642 (0.8%)docs/evals/eu-trial-lay-summary.md, test split
- Wrong numbers left after the check and one repair: 0 / 655docs/evals/eu-trial-lay-summary.md, test split
- Planted errors caught by the code checks (12 held-out trials): 555 / 577 (96.2%)docs/evals/eu-trial-lay-summary.md, test split
- Planted wording errors caught by the grounding judge: 29 / 45 (64%)docs/evals/eu-trial-lay-summary.md, test split
- Flesch-Kincaid grade of the drafts: median (range): 6.1 (4.4-8.0)docs/evals/eu-trial-lay-summary.md, test split
- Numbers in Member-State versions traced to the results cells (12 held-out trials, 6 languages, frozen): 4,002 / 4,205 (95.2%)decosa-api docs/evals/language-pack.md, 27 Sep 2026
- Planted number errors in the translations caught (12 trials, 6 languages): 2,426 / 2,432 (99.8%)decosa-api docs/evals/language-pack.md