Skip to content
decosa

66 · Healthcare · Science and research · live

EU trial lay summary with number grounding

Open the toolJSON

Eval results

Scored on a held-out or test splitRun 27 Sep 2026Eval write-up (decosa-api, access required)

  • Planted errors caught, drafts with citations (code checks)555 / 577 (96.2%)test splitn = 57712 held-out ClinicalTrials.gov trials; numbers, percentages, dates, swapped groups, flipped comparisons, dropped hedges, softened or denied side effects.
  • Correct sentences flagged (false flags), code checks9 / 812 (1.1%)test splitn = 812Adjudicated by the author.
  • Planted errors caught after fixes, fresh trials212 / 217 (97.7%)held outn = 2175 trials kept aside while the test split's misses were fixed; 16 of 347 correct sentences flagged there, from one trial's shortened group names (fixed afterwards).
  • Planted errors caught, citations stripped315 / 577 (54.6%)test splitn = 577
  • Wrong numbers in the model's first drafts5 / 642 (0.8%)test splitn = 642None left after the check and one repair pass (0 / 655).
  • Planted wording errors caught by the grounding judge29 / 45 (64%)test split
  • Flesch-Kincaid grade, median (range)6.1 (4.4-8.0)test splitn = 12
  • Planted errors caught, dev130 / 132dev (tuned on)n = 132The 3 sample trials, used while writing the prompts and the checker.
  • Numbers in Member-State versions traced to the results, 6 languages (frozen)4,002 / 4,205 (95.2%)test splitn = 4,205The 12 test trials' English drafts translated into de, fr, es, it, nl, pl. Most untraced numbers were the checker misreading other languages (fixed afterwards: 4,014 / 4,077, no longer held out).
  • Planted number errors in the translations caught2,426 / 2,432 (99.8%)test splitn = 2,432One digit changed in a sentence both checks had passed; 917 / 920 on the 5 test2 trials. Real errors the check found in the model's translations: 6 Polish sentences that dropped a '0 out of' count (test2) and one Spanish '30,000' left in English format.

Dataset

20 real phase 3 trials with results on ClinicalTrials.gov (19 sponsors, fetched 26 Sep 2026): dev 3, test 12 (run once, frozen), test2 5 (kept fresh for the fixes the test split led to). Drafts by our model; errors planted in code, one per sentence. Member-State versions: the test and test2 drafts translated into six languages on 27 Sep 2026.

Caveats

  • The drafts and the planted errors are ours; no summaries written by people and no published lay summaries were checked.
  • False flags and the 10-summary review were judged by the agent that built the checker; no independent reviewer and no lay-reader test.
  • After the test run its misses were fixed; the test numbers after the fixes (576 / 577) are no longer held out, and the fresh test2 run (212 / 217) checks the fixes.
  • ClinicalTrials.gov records only; CTIS and EudraCT tables were not read.
  • Coverage says which Annex V parts are present, not whether they are adequate; it never says a summary complies.
  • Member-State versions are machine translations with numbers checked in code; the optional meaning check caught all 6 known Polish "0 out of" drops and real errors such as "Vehicle Cream" as a cream for cars, but also raises false alarms (9 of 14 errors on a 240-sentence sample). Grammar is not checked; no native speaker has read them.
  • The fifteen EU languages added on 27 Sep (Qwen3.8-27B route) were run on the 5 test2 trials in five of them (fi, el, hu, ro, ga): 1,604 of 1,624 numbers traced to the results, 1 wrong (Irish); real errors found were a dropped Finnish count and two wrong Irish month names. No native speaker has read any version.

Nightly smoke check

Loading the nightly status…

Result
pass
Run
26 Sep 2026
Latency, this run
n/a
p50 over passed runs
8.2 s
Receipts
3
Model calls
n/a
Tokens
n/a
Cost per run
$0.001

Self-host verification

Verified on 26 Sep 2026: fresh clone, compose up, sample against local model servers

A fresh clone of a decosa-api pre-release build (not yet merged), the api image built from it with DECOSA_LAYSUMMARY_FETCH=0, run against the already-running local Qwen3.8-27B vLLM on the direct route. The rehearsal bundle passed 9 of 9 checks in 0.8 s, the smoke module passed in 0.9 s with 3 attested receipts, a full draft of the ruxolitinib sample took 18.4 s (59 attested receipts, 49 of 49 numbers traced, grade 5.7) and its record verified, and an NCT number was refused with fetching off. Torn down after. Model-server startup itself not re-verified.

Rehearsal bundle: eu-trial-lay-summary.zip (17 KB, 9 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.

Known limits

  • Drafts in English; Member-State versions are machine translations with their numbers checked in code, not their wording.
  • Reads ClinicalTrials.gov records; CTIS and EudraCT results tables are not read.
  • Without citations in a draft, about half of the planted wrong numbers were missed (315 of 577 caught): cite the cells or turn the grounding judge on.
  • The grounding judge can be wrong and is not fully repeatable on the shared gateway; one correct sentence was once called contradicted.
  • Coverage says which Annex V parts are present, never that a summary meets Annex V.

Models and licences

Standard tier. Licence posture: permissive (Apache, MIT or BSD).

  • Results normaliser, number tracing, comparison and hedge checks, side-effect completeness, word lists, Annex V coverage, readability, review file and signed record (no model; CPU)decosa-api lay summary (decosa_api/verticals/laysummary), importing numeric grounding (53) and table cells (57), the grounding judge (17) and the signed record (07)AGPL-3.0-or-later
  • Drafts the six narrative sections with citations, rewrites a failed section once, and judges each sentence against the sourcesQwen3.8-27B (NVFP4)Apache-2.0
  • Member-State versions: each sentence translated, its numbers checked against the English sentence and traced to the results cells in that language: German, French, Spanish, Italian, Dutch, Polish, Portuguese, CzechHy-MT2-7B (the language-pack block)Apache-2.0
  • Member-State versions: the fifteen other EU languages (Swedish, Danish, Finnish, Greek, Romanian, Hungarian, Bulgarian, Croatian, Slovak, Slovenian, Lithuanian, Latvian, Estonian; Irish and Maltese as drafts)Qwen3.8-27B (the language-pack block's route for these languages)Apache-2.0

All quality evidence

Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.

Lite · check a draft in code, no GPU (3)
  • Planted errors caught in drafts with citations (12 held-out trials): 555 / 577 (96.2%)docs/evals/eu-trial-lay-summary.md, test split, 26 Sep 2026
  • Correct sentences flagged (false flags, same drafts): 9 / 812 (1.1%)docs/evals/eu-trial-lay-summary.md, test split, adjudicated by the author
  • Planted errors caught with the citations stripped: 315 / 577 (54.6%)docs/evals/eu-trial-lay-summary.md, test split
Standard · one GPU for the model (hosted demo) (7)
  • Wrong numbers in the model's first drafts (12 held-out trials): 5 / 642 (0.8%)docs/evals/eu-trial-lay-summary.md, test split
  • Wrong numbers left after the check and one repair: 0 / 655docs/evals/eu-trial-lay-summary.md, test split
  • Planted errors caught by the code checks (12 held-out trials): 555 / 577 (96.2%)docs/evals/eu-trial-lay-summary.md, test split
  • Planted wording errors caught by the grounding judge: 29 / 45 (64%)docs/evals/eu-trial-lay-summary.md, test split
  • Flesch-Kincaid grade of the drafts: median (range): 6.1 (4.4-8.0)docs/evals/eu-trial-lay-summary.md, test split
  • Numbers in Member-State versions traced to the results cells (12 held-out trials, 6 languages, frozen): 4,002 / 4,205 (95.2%)decosa-api docs/evals/language-pack.md, 27 Sep 2026
  • Planted number errors in the translations caught (12 trials, 6 languages): 2,426 / 2,432 (99.8%)decosa-api docs/evals/language-pack.md

How we measure · All tools