{"schema_version":"1","site":"https://decosa.ai","id":"eu-trial-lay-summary","num":"66","name":"EU trial lay summary with number grounding","tool_name":"Check a lay summary's numbers","short":"Trial lay summary","blurb":"For sponsors' clinical disclosure teams, medical writers and CROs. Give it a trial with results on ClinicalTrials.gov and it drafts the lay summary the EU Clinical Trials Regulation requires (Article 37(4), content in Annex V), or checks the one you wrote. Every number is traced in code to a results cell, or to arithmetic on the cells the sentence cites, and a wrong one is flagged with the right figure. Comparisons are checked against the table, a result that could be chance has to say so, serious side effects and deaths must be given for every group, and promotional or softening words are flagged. Each claim goes through a grounding judge; Annex V's ten elements and the reading grade are reported. Out come a review file and a signed record of every check. It never says a summary is compliant: the sponsor's medical writer approves.","status":"live","labels":{"industry":["healthcare","science-research"],"job":["draft","review"],"input":["text"],"deploy":["hosted","selfhost"],"status":"live","output":["text","record"],"data":["confidential"],"hardware":"gpu-96","licence":"permissive","mode":"check"},"industries":["healthcare","science-research"],"runs_in":["hosted","selfhost"],"part_of":[],"built_from":["numeric-grounding","grounding","signed-record","language-pack"],"models":"Qwen3.8-27B drafts the sections and judges each claim; number tracing, comparisons, hedges, side-effect checks, coverage and readability are plain code","where":"Hosted or self-host; self-host for results that are not public yet","hardware":"Qwen3.8-27B (1× RTX 5090 32 GB or larger); the number tracing, checks and record run on CPU","final_artifact":"A lay summary in Annex V's order with every number traced to a results cell, the sentences to fix or check, Annex V coverage, the reading grade, a review file and a signed decosa.record.v1.","self_host_first":false,"verification":{"receipt_coverage":"full","summary":"Receipt per model call; every number traced to a results cell id; signed hash-chained record of the sources, the draft and every check","manual_qa":{"hosted":{"date":"2026-09-26","result":"pass","p50_ms":8248,"p95_ms":null,"runs":null,"receipts_per_run":3,"cost_per_run_usd":0.001406},"selfhost":{"date":"2026-09-26","result":"pass","method":"fresh clone, compose up, sample against local model servers","notes":"A fresh clone of a decosa-api pre-release build (not yet merged), the api image built from it with DECOSA_LAYSUMMARY_FETCH=0, run against the already-running local Qwen3.8-27B vLLM on the direct route. The rehearsal bundle passed 9 of 9 checks in 0.8 s, the smoke module passed in 0.9 s with 3 attested receipts, a full draft of the ruxolitinib sample took 18.4 s (59 attested receipts, 49 of 49 numbers traced, grade 5.7) and its record verified, and an NCT number was refused with fetching off. Torn down after. Model-server startup itself not re-verified."},"known_limits":["Drafts in English; Member-State versions are machine translations with their numbers checked in code, not their wording.","Reads ClinicalTrials.gov records; CTIS and EudraCT results tables are not read.","Without citations in a draft, about half of the planted wrong numbers were missed (315 of 577 caught): cite the cells or turn the grounding judge on.","The grounding judge can be wrong and is not fully repeatable on the shared gateway; one correct sentence was once called contradicted.","Coverage says which Annex V parts are present, never that a summary meets Annex V."],"nightly_covers":null},"nightly":"https://api.decosa.ai/verify/status"},"eval_summary":{"metrics":[{"name":"Planted errors caught, drafts with citations (code checks)","value":"555 / 577 (96.2%)","unit":null,"n":577,"split":"test","note":"12 held-out ClinicalTrials.gov trials; numbers, percentages, dates, swapped groups, flipped comparisons, dropped hedges, softened or denied side effects."},{"name":"Correct sentences flagged (false flags), code checks","value":"9 / 812 (1.1%)","unit":null,"n":812,"split":"test","note":"Adjudicated by the author."},{"name":"Planted errors caught after fixes, fresh trials","value":"212 / 217 (97.7%)","unit":null,"n":217,"split":"heldout","note":"5 trials kept aside while the test split's misses were fixed; 16 of 347 correct sentences flagged there, from one trial's shortened group names (fixed afterwards)."},{"name":"Planted errors caught, citations stripped","value":"315 / 577 (54.6%)","unit":null,"n":577,"split":"test","note":null},{"name":"Wrong numbers in the model's first drafts","value":"5 / 642 (0.8%)","unit":null,"n":642,"split":"test","note":"None left after the check and one repair pass (0 / 655)."},{"name":"Planted wording errors caught by the grounding judge","value":"29 / 45 (64%)","unit":null,"n":null,"split":"test","note":null},{"name":"Flesch-Kincaid grade, median (range)","value":"6.1 (4.4-8.0)","unit":null,"n":12,"split":"test","note":null},{"name":"Planted errors caught, dev","value":"130 / 132","unit":null,"n":132,"split":"dev","note":"The 3 sample trials, used while writing the prompts and the checker."},{"name":"Numbers in Member-State versions traced to the results, 6 languages (frozen)","value":"4,002 / 4,205 (95.2%)","unit":null,"n":4205,"split":"test","note":"The 12 test trials' English drafts translated into de, fr, es, it, nl, pl. Most untraced numbers were the checker misreading other languages (fixed afterwards: 4,014 / 4,077, no longer held out)."},{"name":"Planted number errors in the translations caught","value":"2,426 / 2,432 (99.8%)","unit":null,"n":2432,"split":"test","note":"One digit changed in a sentence both checks had passed; 917 / 920 on the 5 test2 trials. Real errors the check found in the model's translations: 6 Polish sentences that dropped a '0 out of' count (test2) and one Spanish '30,000' left in English format."}],"dataset":"20 real phase 3 trials with results on ClinicalTrials.gov (19 sponsors, fetched 26 Sep 2026): dev 3, test 12 (run once, frozen), test2 5 (kept fresh for the fixes the test split led to). Drafts by our model; errors planted in code, one per sentence. Member-State versions: the test and test2 drafts translated into six languages on 27 Sep 2026.","held_out":true,"caveats":["The drafts and the planted errors are ours; no summaries written by people and no published lay summaries were checked.","False flags and the 10-summary review were judged by the agent that built the checker; no independent reviewer and no lay-reader test.","After the test run its misses were fixed; the test numbers after the fixes (576 / 577) are no longer held out, and the fresh test2 run (212 / 217) checks the fixes.","ClinicalTrials.gov records only; CTIS and EudraCT tables were not read.","Coverage says which Annex V parts are present, not whether they are adequate; it never says a summary complies.","Member-State versions are machine translations with numbers checked in code; the optional meaning check caught all 6 known Polish \"0 out of\" drops and real errors such as \"Vehicle Cream\" as a cream for cars, but also raises false alarms (9 of 14 errors on a 240-sentence sample). Grammar is not checked; no native speaker has read them.","The fifteen EU languages added on 27 Sep (Qwen3.8-27B route) were run on the 5 test2 trials in five of them (fi, el, hu, ro, ga): 1,604 of 1,624 numbers traced to the results, 1 wrong (Irish); real errors found were a dropped Finnish count and two wrong Irish month names. No native speaker has read any version."],"date":"2026-09-27","doc_url":"https://decosa.ai/metrics/evals/eu-trial-lay-summary"},"stack":{"summary":"For sponsors' clinical disclosure teams, medical writers and CROs. The EU Clinical Trials Regulation requires a summary of every trial's results for laypersons (Article 37(4); the content is set by Annex V). Give it a trial with results on ClinicalTrials.gov and it drafts that summary, or give it your own draft to check. Code turns the posted results into cells with ids (participant flow, baseline, each outcome with its statistical analysis, adverse events), registry facts and protocol sentences. Qwen3.8-27B drafts the six narrative sections, each sentence citing the ids it used; code fills the identification, sponsor, follow-up and where-to-find-more sections from registry facts and leaves marked gaps for what the registry does not hold. Then every number is traced in code to a results cell, or to arithmetic on the cells a sentence cites (22 of 291 is 8%, about 1 in 13), and a wrong one is flagged with the right figure. Comparisons are checked against the table, a result the posted analysis says could be chance must say so, serious side effects and deaths must be given for every group, and promotional or softening words are flagged. A section that fails a check is rewritten once. Each narrative sentence goes through the grounding judge (vertical 17). It reports Annex V coverage and the reading grade, and writes a review file and a signed record of every check. It never says a summary is compliant: the sponsor's medical writer approves. Member-State versions: the finished English summary is translated sentence by sentence into up to six languages by the language pack (Hy-MT2-7B); each translated number is checked against the English sentence and traced again to the results cells in that language, and the record is signed per language. Domain terms are held to a term bank: the clinical-trials pack (the Clinical Trials Regulation's terms; a vehicle cream is a placebo, not a cream for cars) plus the sponsor's own reviewed entries, with the bank version in each receipt. A native speaker and a lay reader still need to read each version.","tagline":"Draft or check the EU lay summary of a trial's results, with every number traced to the results table.","deployment":"hosted-or-self-host","regulatory_note":"Checked 26 Sep 2026 against the primary sources (links under Tools). Regulation (EU) No 536/2014, Article 37(4): within one year of the end of a trial in all Member States concerned, whatever the outcome, the sponsor submits a summary of the results to the EU database, accompanied by a summary understandable to laypersons whose content is set out in Annex V. Annex V lists ten elements (identification; sponsor name and contact details; where, when, objectives and reasons; the population, including numbers in the Member State concerned, the Union and third countries, age and gender breakdown and eligibility; the investigational medicinal products; adverse reactions and their frequency; overall results; comments on the outcome; whether follow-up trials are foreseen; where to find more). The 6-month deadline for paediatric trials is not in Article 37 itself: the Good Lay Summary Practice (adopted by the Clinical Trials Expert Group, 9 Jul 2021) gives 12 months, 6 for paediatric studies and up to 30 for non-therapeutic phase 1, citing the EU portal specifications (EMA/42176/2014). The expert group's recommendations (v2, 22 Feb 2018) ask for no promotional content, plain language and numeracy, explain that a non-significant difference should be explained to the reader, and call a 6th-grade Flesch-Kincaid level ideal. A cross-sectional study of the 7,547 phase II-IV trials registered in CTIS by November 2025 found that of the 234 legally required to report, 116 (49.6%) fully reported results on time (Bruckner et al., medRxiv preprint, 5 Apr 2026, not peer reviewed). Annex V asks for adverse reactions; the posted tables list adverse events whatever their cause, and the draft says so. This is a drafting and checking aid, not legal or regulatory advice, and it never states that a summary meets Annex V. Model licence: Apache-2.0 (Qwen3.8-27B). ClinicalTrials.gov records are cited by NCT number; results are facts reported by sponsors.","components":[{"id":"laysummary","role":"Results normaliser, number tracing, comparison and hedge checks, side-effect completeness, word lists, Annex V coverage, readability, review file and signed record (no model; CPU)","name":"decosa-api lay summary (decosa_api/verticals/laysummary), importing numeric grounding (53) and table cells (57), the grounding judge (17) and the signed record (07)","hf_repo":null,"license":"AGPL-3.0-or-later","params":null,"quant":null,"vram_gb":0,"memory_gb_estimate":null,"engine":"Python 3.12; Decimal arithmetic; Annex V verbatim from EUR-Lex (read 26 Sep 2026)","receipt_coverage":"partial","in_hosted_demo":null,"tiers":["lite","standard"],"alternative_to":null},{"id":"model","role":"Drafts the six narrative sections with citations, rewrites a failed section once, and judges each sentence against the sources","name":"Qwen3.8-27B (NVFP4)","hf_repo":"nvidia/Qwen3.8-27B-NVFP4","license":"Apache-2.0","params":"27.8B","quant":"NVFP4 (MLP NVFP4, GDN/attention FP8) + FP8 KV cache; MTP head, 3 draft tokens","vram_gb":20,"memory_gb_estimate":null,"engine":"vLLM 0.29.0, temperature 0, thinking off, prefix caching","receipt_coverage":"strong","in_hosted_demo":true,"tiers":["standard"],"alternative_to":null},{"id":"translate","role":"Member-State versions: each sentence translated, its numbers checked against the English sentence and traced to the results cells in that language: German, French, Spanish, Italian, Dutch, Polish, Portuguese, Czech","name":"Hy-MT2-7B (the language-pack block)","hf_repo":"tencent/Hy-MT2-7B","license":"Apache-2.0","params":"7.5B","quant":"BF16","vram_gb":18,"memory_gb_estimate":null,"engine":"vLLM 0.29.0 (Transformers backend, --enforce-eager), gpu-memory-utilization 0.18","receipt_coverage":"partial","in_hosted_demo":true,"tiers":["standard"],"alternative_to":null},{"id":"mt-llm","role":"Member-State versions: the fifteen other EU languages (Swedish, Danish, Finnish, Greek, Romanian, Hungarian, Bulgarian, Croatian, Slovak, Slovenian, Lithuanian, Latvian, Estonian; Irish and Maltese as drafts)","name":"Qwen3.8-27B (the language-pack block's route for these languages)","hf_repo":"Qwen/Qwen3.8-27B","license":"Apache-2.0","params":"27B","quant":"NVFP4","vram_gb":0,"memory_gb_estimate":null,"engine":"The already-served hosted model qwen3.8-27b, through the gateway; no new GPU","receipt_coverage":"strong","in_hosted_demo":null,"tiers":["standard"],"alternative_to":null}],"tiers":[{"id":"lite","label":"Lite · check a draft in code, no GPU","summary":"POST /laysummary/check with grounding off: every number traced, comparisons, hedges, side effects, word lists, Annex V coverage and readability. No drafting and no claim-by-claim grounding.","components":["laysummary"],"hardware":"Any CPU","quality_evidence":[{"metric":"Planted errors caught in drafts with citations (12 held-out trials)","value":"555 / 577 (96.2%)","source":"docs/evals/eu-trial-lay-summary.md, test split, 26 Sep 2026"},{"metric":"Correct sentences flagged (false flags, same drafts)","value":"9 / 812 (1.1%)","source":"docs/evals/eu-trial-lay-summary.md, test split, adjudicated by the author"},{"metric":"Planted errors caught with the citations stripped","value":"315 / 577 (54.6%)","source":"docs/evals/eu-trial-lay-summary.md, test split"}],"latency_note":"no model call; well under a second per draft on CPU (not separately timed)","in_hosted_demo":true,"receipt_coverage":"none","receipt_note":"No model call, so no receipts; the checks are deterministic code.","hosting":null},{"id":"standard","label":"Standard · one GPU for the model (hosted demo)","summary":"Qwen3.8-27B drafts the six narrative sections and judges each sentence; every check and the record are code. This is what the hosted demo runs. Member-State versions: Hy-MT2-7B translates each sentence, and every number is checked again in each language.","components":["laysummary","model","translate"],"hardware":"1x RTX PRO 6000 96 GB (measured) or 1x RTX 5090 32 GB (estimate)","quality_evidence":[{"metric":"Wrong numbers in the model's first drafts (12 held-out trials)","value":"5 / 642 (0.8%)","source":"docs/evals/eu-trial-lay-summary.md, test split"},{"metric":"Wrong numbers left after the check and one repair","value":"0 / 655","source":"docs/evals/eu-trial-lay-summary.md, test split"},{"metric":"Planted errors caught by the code checks (12 held-out trials)","value":"555 / 577 (96.2%)","source":"docs/evals/eu-trial-lay-summary.md, test split"},{"metric":"Planted wording errors caught by the grounding judge","value":"29 / 45 (64%)","source":"docs/evals/eu-trial-lay-summary.md, test split"},{"metric":"Flesch-Kincaid grade of the drafts: median (range)","value":"6.1 (4.4-8.0)","source":"docs/evals/eu-trial-lay-summary.md, test split"},{"metric":"Numbers in Member-State versions traced to the results cells (12 held-out trials, 6 languages, frozen)","value":"4,002 / 4,205 (95.2%)","source":"decosa-api docs/evals/language-pack.md, 27 Sep 2026"},{"metric":"Planted number errors in the translations caught (12 trials, 6 languages)","value":"2,426 / 2,432 (99.8%)","source":"decosa-api docs/evals/language-pack.md"}],"latency_note":"measured: a full draft takes a couple of minutes on the shared gateway, longer under load; a check with grounding takes a call per narrative sentence.","in_hosted_demo":true,"receipt_coverage":"strong","receipt_note":"Every model call is a separate gateway call with a gateway-signed receipt; the signed record lists them all.","hosting":null}],"alternates":[],"services":[{"name":"decosa-api","port":8445,"image":"${DECOSA_REGISTRY}/decosa-api:<tag>","purpose":"GET /laysummary/info, /laysummary/samples; POST /laysummary/sources, /laysummary/draft (SSE or JSON), /laysummary/check; POST /record/verify."},{"name":"vLLM (model)","port":8114,"image":"vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1","purpose":"Qwen3.8-27B NVFP4 behind our gateway (hosted) or called directly (self-host)."}],"tools":[{"name":"Regulation (EU) No 536/2014, Article 37 and Annex V (EUR-Lex)","url":"https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:32014R0536","license":"EU legislation (reuse allowed, Commission Decision 2011/833/EU)","purpose":"The deadline and the ten elements, quoted verbatim in GET /laysummary/info."},{"name":"Summaries of Clinical Trial Results for Laypersons, v2 (EU expert group, 22 Feb 2018), and Good Lay Summary Practice (CTEG, 2021)","url":"https://health.ec.europa.eu/medicinal-products/eudralex/eudralex-volume-10_en","license":"European Commission documents","purpose":"The basis for the readability target (grade 6 ideal, we flag above 8), the neutral-language and numeracy checks, and the paediatric deadline note."},{"name":"ClinicalTrials.gov API v2","url":"https://clinicaltrials.gov/data-api/api","license":"US National Library of Medicine registry; results are facts reported by sponsors","purpose":"The results tables and protocol text, fetched by NCT number (only the number is sent) and cached for a day. Off with DECOSA_LAYSUMMARY_FETCH=0."},{"name":"Bruckner et al., Assessing Compliance with Reporting Requirements in European Phase II-IV Clinical Trials (medRxiv, 5 Apr 2026)","url":"https://www.medrxiv.org/content/10.64898/2026.04.03.26350111v1","license":"preprint (read, not redistributed)","purpose":"The on-time reporting figure in the regulatory note."}],"hardware":[{"tier":"1x RTX PRO 6000 Blackwell 96 GB","fits":true,"notes":"Measured on our server: the hosted demo and the eval ran through the shared gateway."},{"tier":"1x RTX 5090 32 GB","fits":true,"notes":"Estimate: Qwen3.8-27B NVFP4 needs about 20 GB of weights plus KV cache; not run for this tool."},{"tier":"CPU only","fits":true,"notes":"The lite tier (POST /laysummary/check with grounding off: number tracing, comparisons, hedges, side effects, words, coverage, readability) needs no GPU."}],"latency":[{"lane":"a full draft (six sections, repair, grounding), hosted gateway route","typical_ms":115000,"source":"measured on our server 2026-09-26: median over the 12 test trials, 88-289 s, 3 drafts in parallel on a gateway shared with other workloads; the 5 test2 trials took 219-363 s under heavier load"},{"lane":"check a short draft with grounding, hosted gateway route","typical_ms":8248,"source":"measured on our server 2026-09-26: the three-sentence smoke check took 5.0-9.1 s on the shared gateway (3 grounding calls) and 0.9 s self-hosted on the direct route"},{"lane":"check a draft without the model (number tracing, comparisons, side effects, coverage, readability)","typical_ms":null,"source":"no model call; well under a second on CPU (not separately timed)"}],"benchmark":null,"notes":["On 12 held-out trials (577 planted errors in drafts that had passed), the code checks caught 555 (96.2%) with citations kept, and flagged 9 of 812 correct sentences (1.1%). Every miss was a comparison or side-effect wording pattern or a loose number rule; each was fixed with a test, and 5 fresh trials then gave 212 of 217 (97.7%).","The model's first drafts put 5 wrong or untraceable numbers among 642 (0.8%) on the test trials; after the check and one repair pass, none was left.","Without citations (a sponsor's own draft), the number check falls back to the tables of the sentence's section and caught 315 of 577 (54.6%): cite the cells, or turn the grounding judge on.","Every test draft read at Flesch-Kincaid grade 4.4 to 8.0 (median 6.1). Annex V elements 1, 2, 4 and 9 always need the sponsor: the registry does not hold the EU trial number, contact details, participants per Member State or follow-up plans, so the draft marks those gaps instead of inventing them.","The grounding judge caught real slips the number trace cannot: a correct number attached to the wrong outcome, a median called an average, 'took' for 'analysed'. It also refuses some plain definitions (a placebo is a dummy treatment); a glossary source was added for that after the eval.","Everything was measured on ClinicalTrials.gov records with drafts from our own model. It has not been compared with published lay summaries or tested with lay readers."]},"buyer_facts":[{"label":"Data retention","value":"Nothing stored: the study record, the draft and the result live in memory for the request, and logs carry counts only. A registry record fetched by NCT number is cached for a day. The signed record holds hashes, statuses and receipt ids, and the text of each sentence."},{"label":"What leaves the box","value":"Hosted: every model call goes through our gateway to the GPU serving Qwen3.8-27B, and an NCT number you type is sent to ClinicalTrials.gov. Self-hosted with DECOSA_LAYSUMMARY_FETCH=0 on the direct route: nothing leaves the box."},{"label":"What it will not say","value":"That a summary is compliant or meets Annex V. It says what it checked and marks what the registry does not hold (EU trial number, contact details, participants per country, follow-up plans) for the sponsor to add."},{"label":"Input formats","value":"A ClinicalTrials.gov NCT number or API v2 study record (CTIS and EudraCT tables have to be pasted in that shape), an optional protocol synopsis (20,000 characters), and for checking, a Markdown draft of up to 16,000 characters in English, German, French, Italian, Dutch or Spanish. Member-State versions: up to six of de, fr, es, it, nl, pl, pt, cs per request."},{"label":"Typical run","value":"A full draft: dozens of model calls, a few cents at the gateway list price, a minute or more on the shared gateway (faster self-hosted on the direct route). A short check with grounding: a few calls, a fraction of a cent. Member-State versions with the optional meaning check: two more calls per sentence per language, which adds minutes for many sentences in several languages on the busy gateway."}],"hosted_now":{"needs":["qwen3.8-27b","translation"],"off":["translation"],"live_by_default":false,"live_status":"https://api.decosa.ai/status"},"data_handling":{"page":"/data#eu-trial-lay-summary","self_host":{"level":"confidential","leaves":"identifiers","summary":"Runs on your machine; by default only short identifiers or a digest go to the public services listed in external_calls."},"hosted":{"level":"operator-processed","demo_only":false,"summary":"TLS to Decosa's server, then decrypted and processed by Decosa's API server, with the open models run by NEAR AI through OpenRouter, with Reka AI as the only fallback under Decosa's account.","gpus":"operator-contracted","third_parties":[],"retention":"Nothing stored: the study record, the draft and the result live in memory for the request, and logs carry counts only. A registry record fetched by NCT number is cached for a day. The signed record holds hashes, statuses and receipt ids, and the text of each sentence.","used_for_training":false,"encrypted_while_processed":false},"sealed_tier":{"applies":false,"note":"The sealed tier (raw chat only, never use-case pipelines) is paused at launch (/docs/sealed-tier)."},"external_calls":[{"to":"ClinicalTrials.gov (US National Library of Medicine)","route":"both","sends":"identifiers","what":"The NCT number you ask it to fetch. Never your draft, your inputs or unpublished results.","default":"on","off":"Set DECOSA_LAYSUMMARY_FETCH=0, or send the study record yourself as \"study\"."}]},"console":{"href":"/tools/life-sciences/eu-trial-lay-summary","input":"laysummary","lanes":[{"id":"draft","title":"Lay summary, in Annex V's order","kind":"markdown"},{"id":"numbers","title":"Every number traced to a results cell","kind":"list"},{"id":"checks","title":"Comparisons, hedges, side effects, wording, grounding","kind":"list"},{"id":"record","title":"Coverage, readability, review file and signed record","kind":"json"}],"samples":[{"n":1,"id":"icodec-weekly-insulin","title":"Icodec weekly insulin","deep_link":"/tools/life-sciences/eu-trial-lay-summary?sample=1&autorun=0"},{"n":2,"id":"ruxolitinib-covid","title":"Ruxolitinib covid","deep_link":"/tools/life-sciences/eu-trial-lay-summary?sample=2&autorun=0"},{"n":3,"id":"delgocitinib-teens","title":"Delgocitinib teens","deep_link":"/tools/life-sciences/eu-trial-lay-summary?sample=3&autorun=0"},{"n":4,"id":"icodec-check","title":"Icodec check","deep_link":"/tools/life-sciences/eu-trial-lay-summary?sample=4&autorun=0"}],"deep_link_params":{"sample":"1-based index into samples, or a sample id","autorun":"1 = start the run once the sample is loaded; 0 (default) = only preselect","reduce-motion":"1 = turn off animations"}},"api":{"base":"https://api.decosa.ai","contract":"/api/contract.json","contract_markdown":"/api/contract.md","reference":"/docs/api","keys":"/account/keys"},"prompts":{"hosted":"/prompts/eu-trial-lay-summary-hosted.md","selfhost":"/prompts/eu-trial-lay-summary-selfhost.md","assemble":"/prompts/eu-trial-lay-summary-assemble.md","mac":"/prompts/eu-trial-lay-summary-mac.md"},"rehearsal":{"bundle":"/samples/eu-trial-lay-summary.zip","bundle_url":"https://decosa.ai/samples/eu-trial-lay-summary.zip","folder":"/samples/eu-trial-lay-summary/","expected":"/samples/eu-trial-lay-summary/expected.json","files":["/samples/eu-trial-lay-summary/expected.json","/samples/eu-trial-lay-summary/inputs/draft.md","/samples/eu-trial-lay-summary/inputs/study.json"],"bytes":17111,"checks":["the results are read into cells with no model call, and the icodec serious side-effect count is cell C114 (22 of 291)","the correct glargine sentence is traced to its cell and passes","the grounding judge supports it","the wrong icodec count (23) is flagged as a mismatch","with the table's own figure as the nearest","'No one in the trial died' is flagged against the table","the signed record verifies","a record whose draft hash was changed no longer verifies","every model call has a signed receipt"],"licence":"Study record: ClinicalTrials.gov NCT04880850 (US National Library of Medicine registry; results are facts reported by the sponsor, reproduced unchanged). Draft: written for this bundle (CC0). Part of decosa-api, which will be released under AGPL-3.0-or-later; until then the source is on request.","about":"The ONWARDS 4 trial (NCT04880850, insulin icodec weekly vs insulin glargine daily) as posted on ClinicalTrials.gov, and a three-sentence side-effects section: the glargine count is right (25 of 291), the icodec count is wrong (23; the table says 22), and 'No one in the trial died' is false (the table has 2 and 1 deaths). The check must trace the right sentence to its cell and have the grounding judge support it, flag the wrong count with cell C114 as the nearest figure, flag the death claim against the table, and sign a record that verifies and fails once changed. Sources first come back with no model call.","run":{"containers":"docker compose exec api python scripts/rehearse.py eu-trial-lay-summary","checkout":"python scripts/rehearse.py eu-trial-lay-summary --bundle eu-trial-lay-summary.zip --base-url http://127.0.0.1:8445","mac":".venv/bin/python scripts/rehearse.py eu-trial-lay-summary"},"guidance":"Set up with a coding agent (we recommend Claude Code with Claude Opus 5.5; any capable coding agent works) on mock data only, run the rehearsal until every check passes, then run your own data locally yourself. Never give the agent real data during setup."},"hardware_fit":{"check":"/self-host/hardware?use=eu-trial-lay-summary","data":"/api/hardware.json","tiers":[{"id":"lite","gpu_gb":0,"basis":null,"unknown":[]},{"id":"standard","gpu_gb":75.6,"basis":"stack","unknown":[]}],"mac":{"fit":"partial","memory_gb":32}},"links":{"page":"/tools/life-sciences/eu-trial-lay-summary","json":"/use-cases/eu-trial-lay-summary.json","metrics":"/metrics/eu-trial-lay-summary","console":"/tools/life-sciences/eu-trial-lay-summary","console_sample":"/tools/life-sciences/eu-trial-lay-summary?sample=1&autorun=0","stack":"/tools/life-sciences/eu-trial-lay-summary#stack","try_live":"/tools/life-sciences/eu-trial-lay-summary","watch":"/tools/life-sciences/eu-trial-lay-summary","build":"/tools/life-sciences/eu-trial-lay-summary#build","self_host":"/tools/life-sciences/eu-trial-lay-summary#self-host","prompts":{"hosted":"/prompts/eu-trial-lay-summary-hosted.md","selfhost":"/prompts/eu-trial-lay-summary-selfhost.md","assemble":"/prompts/eu-trial-lay-summary-assemble.md","mac":"/prompts/eu-trial-lay-summary-mac.md"}}}