71 · Science and research · Healthcare · live
CSR number-to-table verifier
Eval results
Scored on a held-out or test splitRun 27 Sep 2026Eval write-up (decosa-api, access required)
- Planted number errors caught92 / 96test splitn = 96held-out synthetic CSRs, 8 errors each: transposition, other arm, wrong N, wrong denominator, rounding, stale cut; re-measured 30 Sep 2026
- Planted number errors caught, second writer93 / 96test splitn = 96the held-out CSRs with each paragraph rewritten by Qwen3.8 (numbers kept)
- Planted number errors caught, ClinicalTrials.gov tables170 / 178held outn = 17830 trials' posted results as CSR tables, template narrative
- False flags per 100 numbers on clean reports0 (0 / 899)test splitn = 899
- Traced numbers citing the gold cell99.1%test splitn = 884
- Time per 100 pages129 s (JSON); 300 s (PDF); 493 s (scan)test split
Dataset
Synthetic CSRs of fictional phase 3 trials (sections 10-12, 11 tables each, an earlier data cut): dev seeds 1-4, held-out seeds 101-112 with a template narrative and again rewritten by Qwen3.8; 30 ClinicalTrials.gov trials' posted results laid out as CSR tables with a template narrative. Each report scored clean and with 8 planted errors.
Caveats
- Same author wrote the synthetic reports, the planter and the prompts; real CSRs are longer and messier.
- The ClinicalTrials.gov narratives come from templates, not from a sponsor's CSR.
- No real CSR and no comparison with a QC reviewer's findings.
- Misses are mostly the other arm's value when nothing else in the sentence breaks (3 of 24 held out). Numbers inside real arm names (for example 'Chondroitin 4&6') cause most false flags on the ClinicalTrials.gov set.
Nightly smoke check
Loading the nightly status…
- Result
- pass
- Run
- 27 Sep 2026
- Latency, this run
- n/a
- p50 over passed runs
- 6.5 s
- Receipts
- 6
- Model calls
- n/a
- Tokens
- n/a
- Cost per run
- $0.004
Self-host verification
Verified on 27 Sep 2026: fresh clone into a clean directory, api image from docker/api/Dockerfile, compose api with a named volume, direct route to the local Qwen3.8-27B, the running document reader (:8497) and retrieval (:8499) services, local signing; torn down after
The rehearsal bundle passed 15/15 in 62 s; every receipt attested. The reader and retrieval services were the running ones, not built from the compose file here.
Rehearsal bundle: csr-number-verifier.zip (15 KB, 15 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.
Known limits
- Hosted verification ran on the pre-release server (decosa-api the pre-release branch on our server, gateway route). The production API gets this tool when the branch merges.
- Measured on synthetic CSRs written by the same author as the prompts and on ClinicalTrials.gov results with a template narrative; not on a real sponsor CSR or against a QC reviewer's findings.
- Figures, listings and patient narratives are not read; RTF and SAS outputs must be exported to PDF or rows first.
- A PDF run reads up to 40 pages; split a longer report by section.
- An unflagged number is not proven right: 'matched by value only' means the value was found in one cell, not that the sentence was matched to it. A run where model calls failed is marked incomplete.
Models and licences
Standard tier. Licence posture: permissive (Apache, MIT or BSD).
- Finds the numbers, keeps the tables as typed cells, compares and recomputes in code, names the likely slip, checks in-text tables against their TLF, writes the QC report and the signed record (no model; CPU)decosa-api CSR verifier (decosa_api/verticals/csr), importing the document reader, the evidence retrieval block and the numeric-grounding block's rounding helpersAGPL-3.0-or-later
- One call per paragraph: which cell each number claims to report (by arm, row and timepoint), or which cells a derived number is computed from; a second look for numbers left unplaced; a second reading (values hidden) between two neighbouring cellsQwen3.8-27B (NVFP4)Apache-2.0
- Finds the tables a paragraph most likely reports in a long report (after the tables it names), with a signed receipt per searchEvidence retrieval block: Qwen3-Embedding-0.6B + Qwen3-Reranker-4B (decosa-retrieval service)Apache-2.0
- PDFs and scans: finds and orders the regions of each page (Docling Heron layout) and reads tables as cells with spans (PaddleOCR-VL-1.6); born-digital text comes from the PDF's text layerDocument reader block: Docling 2.130 (Heron layout) + PaddleOCR-VL-1.6 (0.9B)Apache-2.0 (PaddleOCR-VL-1.6 weights, Heron layout weights); MIT (Docling)
All quality evidence
Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.
Lite · text and rows, keyword search (1)
- Planted number errors caught, held-out synthetic CSRs, keyword search: 93 / 96 (96.9%), 1 false flag in 899 clean numbersdecosa-api docs/evals/csr-number-verifier.md, test split, BM25 ablation
Standard · retrieval + document reader + the model (hosted demo) (6)
- Planted number errors caught, held-out synthetic CSRs (template narrative): 92 / 96 (95.8%)decosa-api docs/evals/csr-number-verifier.md, test split
- Planted number errors caught, held-out synthetic CSRs rewritten by a second writer (Qwen paraphrase, numbers kept): 93 / 96 (96.9%)decosa-api docs/evals/csr-number-verifier.md, test-para split
- Planted number errors caught, 30 ClinicalTrials.gov trials laid out as CSR tables: 170 / 178 (95.5%)decosa-api docs/evals/csr-number-verifier.md, ctgov split
- False flags on clean reports (per 100 numbers checked): 0 per 100 (0 / 899) on held-out synthetic; 0.11 (1 / 899) rewritten; 2.04 (19 / 931) on ClinicalTrials.gov tablesdecosa-api docs/evals/csr-number-verifier.md
- Traced numbers citing the right cell: 99.1% (884 numbers, held-out synthetic); 98.2% (901, ClinicalTrials.gov)decosa-api docs/evals/csr-number-verifier.md
- Time per 100 pages: 129 s from JSON; 300 s from a born-digital PDF, 493 s from a scan (reading included), shared gatewaydecosa-api docs/evals/csr-number-verifier.md