CSR number-to-table verifier (71): eval
27 Sep 2026. On a pre-release build. Runners: scripts/csr_eval.py (JSON path), scripts/csr_eval_pdf.py
(PDF path), scripts/csr_ctgov.py (public set). Scores, misses and false flags per set: docs/evals/csr-number-verifier/.
Model: Qwen3.8-27B through the model gateway (every call receipted), measured under the shared gateway's load.
What is measured
A CSR narrative (sections 10 to 12) and its tables go in; every number comes back traced to a cell, recomputed, or flagged. Each trial is scored twice: clean, and with 8 planted errors (at most one per sentence), cycling through the six kinds medical-writing QC looks for:
| Kind | What the planter does |
|---|---|
| transposition | swaps two different digits (57.5 -> 75.5) |
| other_arm | writes the value of the same row in another arm's column |
| wrong_n | changes a population N (randomised, analysis set, column N) by 2 to 9 |
| denominator | recomputes a percentage on another population's N (randomised instead of FAS, and so on) |
| rounding | changes the last printed digit by one (71.9% -> 71.8%) |
| stale | writes the value of the same cell at the earlier (interim) data cut |
- Catch rate: a planted error is caught when a flag (mismatch, untraceable, or a derived number that doesn't compute) lands on the planted number. A flag on another number in the same sentence (a percentage that no longer computes from a planted N) is counted separately as a knock-on, not as a catch.
- False flags: every flag on a clean document.
- Citation accuracy: traced numbers on clean documents whose cited cell is the gold cell (an in-text table that copies the gold TLF, same cell, counts as right; the stricter "gold table exactly" is also given); and caught mismatches that cite the gold cell.
- Time per 100 pages: wall time for the whole run divided by pages (the JSON path estimates pages as the sample PDFs lay them out; the PDF path counts real pages).
Sets
All built before any held-out run; seeds are fixed in the runners. Licence of every input: synthetic sets are invented
by decosa_api/verticals/csr/synth.py (fictional drugs, sponsors and trials; AGPL-3.0-or-later with decosa-api). The public set
uses ClinicalTrials.gov results (terms at https://clinicaltrials.gov/about-site/terms-conditions, last updated 31 January
2023: attribute ClinicalTrials.gov, show the processing date, state modifications). Fetched 27 Sep 2026. Modifications:
2-3 arm completed phase 3 trials chosen by condition; participant flow, baseline, the first primary outcome, adverse-event
totals and the most frequent non-serious AEs laid out as CSR tables; percentages computed from the posted counts; a
template narrative written from them. EMA clinical data (Policy 0070) is not used: its terms of use allow general
information and non-commercial research use only (EMA/144064/2019, 21 March 2019, Annex 1 and 2), which rules out a
commercial demo.
| Set | Trials | Documents | Planted | Numbers on clean docs | Role |
|---|---|---|---|---|---|
| dev | 4 synthetic (seeds 1-4) | 8 | 32 | 304 | used to fix prompts, context rules and code |
| test | 12 synthetic (seeds 101-112) | 24 | 96 | 899 | held out |
| test-para | the test trials, each paragraph rewritten by Qwen3.8 (numbers kept; 113 of 129 paragraphs passed the number check, the rest kept their template text) | 24 | 96 | 899 | held out: a different writer |
| ctgov | 30 ClinicalTrials.gov trials, 20 conditions | 60 | 178 | 931 | held out: real results tables (no earlier cut, so no stale plants) |
| 6 test trials printed as PDFs (and as scans) | 12 + 12 | 48 + 48 | held out: the document reader path |
Results
Product configuration: values shown to the pointer, second reading on, hybrid retrieval.
| Set | Caught | Catch rate | False flags on clean docs | per 100 numbers | Citation (traced) | Citation (mismatch) | s per 100 pages |
|---|---|---|---|---|---|---|---|
| dev | 32/32 | 100% | 0 on 4 docs | 0.00 | 100.0% (298) | 92.9% (28) | 77 |
| test | 92/96 | 95.8% | 0 on 12 docs | 0.00 | 99.1% (884) | 91.8% (85) | 82 |
| test-para | 93/96 | 96.9% | 1 on 12 docs | 0.11 | 97.7% (879) | 86.7% (83) | 205 |
| ctgov | 170/178 | 95.5% | 19 on 30 docs | 2.04 | 98.2% (901) | 96.8% (157) | 193 |
By kind (test / test-para / ctgov): transposition 24/24, 24/24, 53/53; wrong N 12/12, 12/12, 18/18; denominator 12/12, 12/12, 24/24; rounding 12/12, 12/12, 29/29; stale 12/12, 12/12, (none); other arm 20/24, 21/24, 46/54.
Re-measured 30 Sep 2026 (dev and test rows above). After the model server was restarted on 28 Sep (image and video
input turned on), dev gave 4 false flags on one clean report (mar-393-s4), on the old code too: a sentence citing Table
14.3.2 for numbers in Table 14.3.1.1, and the pointer now paired the right row and column with the named table's key
("T11.R4C3" where T11 has three rows). A cell that does not exist used to count as placed, so the number came back
untraceable. Now it is unplaced and gets the second look, which is told which cells did not exist (tests/test_csr.py).
Measured on our server's direct route (the same weights and server as the gateway, not through it): dev 32/32 with 0
false flags; test 92/96 (other arm 20/24), 0 false flags in 899 clean numbers, 99.1% of 884 traced numbers citing the
gold cell, 82 s per 100 pages. The one test miss that is new (zen-368-s106, the other arm's 51.0%) was caught in a repeat
run of that report with the new rule and with the old one: run-to-run variation of the served model, not the change.
test-para, ctgov and the PDF path were not re-run. The 27 Sep score files are kept as score-{dev,test}-llm-0927.json.
Citation "gold table exactly" (not counting in-text copies): test 88.2%, test-para 86.6%, ctgov 98.2%. Most of the difference is the narrative citing Table 12-1 (the in-text copy) where the gold says 14.3.1.1.
PDF path (6 held-out test trials, 12 pages each, clean and planted, read by the document reader service):
| Form | Caught | False flags on 6 clean docs | per 100 numbers | Tables found and named | Reading s / 100 pages | Total s / 100 pages |
|---|---|---|---|---|---|---|
| born-digital PDF | 47/48 | 4 | 0.89 | 128/132 | 93 | 300 |
| scan (image-only, tilted, speckled) | 48/48 | 0 | 0.00 | 132/132 | 254 | 493 |
(The PDF path counts a catch when a flag in the planted number's section names that number, or a derived number in that section does not compute; its paragraph ids differ from the JSON path's.)
The code tier alone (gold cells given, no model)
The same sets with the gold cell handed to the code: test 96/96 and 0 false flags; test-para 95/96 and 3; ctgov 176/178 and 34. So the code comparison and slip rules are not what limits catches; the pointer is. The ctgov false flags in this row come from numbers in arm and outcome names (see below) that no table holds.
Ablations on test
| Configuration | Caught | False flags / 100 numbers | Citation (traced) | s / 100 pages |
|---|---|---|---|---|
| product (values shown, second reading, hybrid retrieval) | 93/96 | 0.00 | 99.1% | 129 |
| no second reading | 92/96 | 0.00 | 99.5% | 109 |
| BM25 retrieval only (no embedder or reranker) | 93/96 | 0.11 | 100% | 79 |
| values hidden from the pointer | 96/96 | 1.89 | 95.2% | 150 |
- Hiding the values makes the pointer place numbers by meaning, so it catches every other-arm slip; but it loses its
footing where a sentence cites one table for numbers printed in another (the SAE sentence citing the by-term table),
and false flags rise from 0 to 17 on 12 clean documents. Kept as an option (
DECOSA_CSR_BLIND=1), off by default. - BM25 alone did as well as hybrid retrieval here: these reports have 11 tables, and the tables a paragraph names are put first anyway. Retrieval should matter on full CSRs with a hundred or more TLFs; not measured.
Where it fails
- The other arm's value, when the pointer follows the value. 3 of 24 (test), 8 of 54 (ctgov). The pointer is told not to choose a cell by value, but when the written number sits in the neighbouring arm's cell it sometimes points there. Most such slips are still caught through their sentence (a percentage that no longer computes from the count); the counted misses are the ones where nothing else in the sentence breaks, like "arthralgia (8.6% and 8.6%)". The values-hidden option catches all of them at a cost in false flags.
- Numbers inside names. Real arm and outcome names carry numbers ("Chondroitin 4&6 Sulfate", "Group 1/2", "SYNC TMS"); the context rules catch doses, weeks, CI levels, thresholds and acronym names (PASI 75, IGA 0/1) but not these. Most of the 19 ctgov false flags are two trials with such names. A medical writer can clear them in a glance, but they are noise.
- A sentence citing one table for numbers in another could leave numbers untraceable (seen on dev); since 30 Sep a cell that does not exist gets a second look, and the dev report that showed it is clean.
- Wrong N and stale values get a generic slip label when the pointer cites an in-text copy; fixed for stale during the build (the earlier cut of the copy's source TLF is now consulted), not re-measured.
- One test-para miss is an eval artefact: the paraphrase relocation put a planted value on the "12" of "Table 12-1".
- The document reader dropped a "Table 14.x" number line on some born-digital pages (the table is then kept under a page-based id and found by retrieval, but a sentence naming that table no longer pins it): 128 of 132 tables named on PDFs, all 132 on scans.
- Synthetic narratives come from one author's templates (and a Qwen rewrite of them); the CT.gov narratives are templates over real tables. Not measured on a real sponsor CSR (none with a commercial-use licence was available).
Time and cost
- JSON path: 129 s per 100 pages on test (about 12 pointer calls per 12-page report, 4 in parallel), 193-205 s on test-para and ctgov under a loaded gateway. About 18k prompt and 1.8k generated tokens per 12-page report: roughly 140k prompt and 14k generated tokens per 100 pages, about $0.06 at the gateway list price ($0.30 / $1.50 per million).
- PDF path: reading adds 93 s per 100 pages (born-digital) or 254 s (scans); 300 and 493 s per 100 pages end to end.
- Smoke (sections 11.4 and 12 of the sample): about 10 s, 6 receipted calls, $0.005.
Checkable properties of the sample run (rehearsal/csr-number-verifier)
- The six seeded numbers are flagged in their sections: 75.5 (11.2), 71.8% (10.1), 313 (11.1), 53.5% (12.2.1), 2.2% (12.3.1); the other-arm 72 in 11.4.3 is caught through its percentage (derived_wrong).
- The stale cell in in-text Table 11-1 is found against its source TLF 14.2.1.
- The clean report comes back with at most one flag (0 in the recorded runs).
- The signed record verifies, and fails once its flag count is changed.
- The PDF of the same report is read into at least ten of its eleven tables, and at least four seeded numbers are flagged.
- Every model call has a signed receipt.
Hosted (pre-release server, gateway route) and self-hosted (fresh clone, compose, direct route to the local Qwen3.8, running reader and retrieval services) rehearsals both pass 15/15.
Tuning on the test sets
None. Prompts, context rules and code were changed only while looking at dev. Two generator fixes were made before the held-out sets were built (an ambiguous deaths sentence, two sentences missing "respectively") and the sets were rebuilt; one classification fix (stale via an in-text copy) came after the test runs, from a demo run, and affects slip labels only. The values-hidden pointer was tried on dev, looked worse on false flags, and was run on test as an ablation.
Verdict
A head of medical writing would pay for this as a QC pass after authoring, not as a replacement for the QC reviewer. On held-out synthetic and public-table reports it caught 95-97% of seeded number errors, with no false flags on clean synthetic reports and about 2 per 100 numbers on real ClinicalTrials.gov tables, citing the right cell 98-99% of the time, for about $0.06 and 2-8 minutes per 100 pages. What's missing before a sponsor would rely on it: a run on a real CSR with its real TLFs (and the RTF outputs they come in), full-size TLF packages (hundreds of tables), handling of numbers in names, figures and listings, and GxP validation if it is ever used inside a validated process.