Document reader (block): eval
Date: 26 Sep 2026; production runtime on our server measured 27 Sep 2026 (section 5). Block: decosa_api.docreader +
services/docreader. Lab code: scripts/docreader_lab/. Every number below is in
docs/evals/document-reader/{ab-dev,test,bench,hcc-dev,hcc-test,hcc-fresh,hcc-fresh2,our server}.json.
What was built and what was measured
Pipeline: Docling (MIT) finds and orders the regions of a page (layout model docling-project/docling-layout-heron
@8f39ad3, Apache-2.0); a ~1B page parser reads each region (text, or a table as HTML with spans); Qwen3.8-27B re-reads a
region when the parser's mean token probability is under 0.90 and reads form fields tied to element ids. Born-digital
PDF pages take their text from the PDF's own text layer; only tables are read from pixels, and every number in them is
checked against that layer.
The parser was chosen by an A/B on dev slices, then only the winner ran on held-out slices, once. All four candidates got the same Docling regions and their own documented element-level prompts, greedy decoding, bf16, on the Mac Studio (M3 Ultra). No prompt or threshold was tuned on a test slice. The one pipeline bug found on dev (reading order, below) was fixed before any test run. The HCC test split and the first fresh split each found one reading bug; each was fixed afterwards and measured on the next untouched split (section 4).
Licences (checked 26 Sep 2026)
| Candidate | Revision | Licence | Note |
|---|---|---|---|
| PaddleOCR-VL-1.6 (chosen) | PaddlePaddle/PaddleOCR-VL-1.6@c5630ab |
Apache-2.0, LICENSE file in the repo | 0.9B |
| MinerU2.5-Pro-2605 | opendatalab/MinerU2.5-Pro-2605-1.2B@bff20d4 |
Apache-2.0 (card) | Trap: the older MinerU2.5-2509-1.2B card now says AGPL-3.0; the MinerU application repo is not Apache either (GitHub: NOASSERTION). mineru-vl-utils is Apache-2.0. Pin the Pro-2605 weights. |
| TeleOCR | StarDoc-AI/TeleOCR@a614331 |
Apache-2.0 per the card metadata; no LICENSE file in the repo | remote code written for transformers 4.x (needs 4.57) |
| OvisOCR2 | ATH-MaaS/OvisOCR2@1fc9221 |
Apache-2.0, LICENSE file | 0.8B |
| Docling + layout | docling 2.130.0; docling-layout-heron@8f39ad3 |
MIT + Apache-2.0 | |
| Datalab Marker / Surya / Chandra | — | Out: the weights forbid use that competes with Datalab's API | not tested |
Datasets (evaluation only; nothing redistributed, no images in this repo):
| Set | Licence | Used |
|---|---|---|
| OmniDocBench (1,651 pages) | repo Apache-2.0; the HF dataset card has no licence field; pages come from many publishers | English pages only: dev 30, test 84 (stratified by source, seed 48) |
| FinTabNet (OTSL copy, test shard 0) | card says "other"; IBM released FinTabNet under CDLA-Permissive [U: not re-read today] | dev 40 tables, test 200 (196 of them have numeric cells) |
| IAM handwriting lines (Teklia copy) | the copy is tagged MIT; the IAM database's own terms are non-commercial research [U: the IAM site did not answer today] | dev 40, test 200 lines |
| CORD-v2 receipts | CC-BY-4.0 | dev 25, test 75 |
| FUNSD forms | research use per its authors [U: the licence text did not render on its site today] | dev 15, test 35 |
IAM and FUNSD are fine for measuring; keep them out of anything we ship or sell.
1. Parser A/B (dev slices)
| OmniDocBench text NED ↓ | OmniDocBench word F1 | FinTabNet TEDS | FinTabNet numbers tied, cell for cell | IAM CER | |
|---|---|---|---|---|---|
| PaddleOCR-VL-1.6 | 0.228 | 0.830 | 0.963 | 98.3% (1,015 / 1,033) | 4.9% |
| MinerU2.5-Pro-2605 | 0.235 | 0.805 | 0.973 | 87.1% | 4.0% |
| TeleOCR | 0.207 | 0.850 | 0.895 | 82.2% | 12.0% |
| OvisOCR2 | 0.236 | 0.845 | 0.795 | 61.4% | 3.6% |
| Tesseract 5.3.4 (today) | 0.333 | 0.703 | no tables | — | 42.0% |
Speed and memory, solo (one parser at a time, nothing else on the GPU, 6 pages / 50 regions, Mac Studio M3 Ultra):
| seconds per page (reading only) | tokens/s | peak memory | runtime | |
|---|---|---|---|---|
| PaddleOCR-VL-1.6 | 3.3 | 258 | 2.9 GB (MLX peak) | mlx-vlm 0.7.3 |
| MinerU2.5-Pro-2605 | 11.6 | 71 | 5.7 GB (MPS) | transformers 5.17 |
| OvisOCR2 | 22.5 | 44 | 6.8 GB (MPS) | transformers 5.17 |
| TeleOCR | 23.7 | 43 | 7.5 GB (MPS) | transformers 4.57 |
Plus Docling's layout model: about 0.15 to 1 s per page and under 1 GB. The runtimes differ: PaddleOCR-VL under transformers on MPS ran at about 7 tokens/s (an op falls back), so it runs on MLX; on the 5 dev pages both runtimes read alike (text NED 0.092 vs 0.093).
Choice: PaddleOCR-VL-1.6. The table numbers decide it: numeric grounding will quote numbers read from tables, and Paddle tied 98.3% of numeric cells on dev against 87% for the next. Its page text is within 0.02 NED of the best, it is 3.5 to 7 times faster on the Mac at half the memory, and its handwriting gap is what the Qwen3.8 re-read covers. TeleOCR reads page text best but tied 82% of table numbers and has 12% CER on handwriting; MinerU has the best table structure (TEDS) but places numbers in the wrong cell more often.
Bug found on dev: the first layout step took Docling's assembled-element order, which interleaves columns. Taking the
order of Docling's document items fixed it (page NED 0.38 → 0.21 to 0.24 for every parser). The parsers' readings did not
change (same regions), so the dev runs were re-sequenced, not re-run (scripts/docreader_lab/reorder_runs.py).
2. Held-out (PaddleOCR-VL-1.6, run once)
| Result | Tesseract on the same slice | |
|---|---|---|
| OmniDocBench, 84 English pages: text NED | 0.148 | 0.239 |
| OmniDocBench: word F1 | 0.876 | 0.763 |
| OmniDocBench: 24 tables, TEDS / TEDS-S | 0.818 / 0.858 | none |
| FinTabNet, 200 tables: TEDS / TEDS-S | 0.930 / 0.944 | none |
| FinTabNet: numeric cells tied, cell for cell | 83.8% (4,488 / 5,353) | none |
| FinTabNet: numbers read correctly anywhere in the table | 94.7% | none |
| FinTabNet: tables with every number tied | 147 / 196 | none |
| IAM, 200 handwritten lines: CER (Paddle alone) | 4.7% | 52.9% |
| IAM: Qwen3.8-27B alone | 3.6% | |
| IAM: the block (Qwen re-reads when Paddle's mean token probability < 0.90) | 4.0%, 81 of 200 lines re-read | |
| Seconds per page on the Mac (shared GPU during the run) | 6.0 | 2.3 |
Numbers from tables are not yet trustworthy on their own. Page 48 set a kill bar: 99% of table numbers must tie out on FinTabNet before numeric grounding trusts them. Held out, Paddle ties 83.8% (dev looked better at 98.3% on 40 tables; the test slice is 5 times larger). About a third of the misses are misread digits (94.7% are read correctly somewhere in the table) and the rest are cells in the wrong row or column. So:
- born-digital PDFs: every table number is checked against the PDF's own text layer (
numbers.in_text_layer); a number that is not there is flagged; - scans: table numbers carry the parser's confidence and must be confirmed by a person before numeric grounding uses them. The block does not claim tie-out on scans.
3. Forms (Qwen3.8-27B, the block's forms prompt)
The block's method sends the page image plus the parser's elements (id: text) and returns fields tied to element ids. The comparison is the same prompt with the image alone (no elements, so no boxes).
| dev | test | |
|---|---|---|
| CORD receipts, 9 total fields, field F1: block (image + elements) | 0.854 | 0.875 |
| CORD: image only | 0.877 | 0.895 |
| FUNSD forms, key F1 / value F1: block | 0.747 / 0.621 | 0.586 / 0.615 |
| FUNSD: image only | 0.766 / 0.615 | 0.501 / 0.489 |
| Filled values found in the text of the element they cite ("grounded") | CORD 156 / 325; FUNSD 323 / 410 |
On short receipts the element list costs about 2 points of F1 against reading the image alone; on dense forms it helps
(value F1 0.615 vs 0.489). The block keeps the element list because it is what ties each value to a page and box. A value
that is not in its element's text is returned with grounded: false: on CORD that is about half the values (receipt
prices are often split across regions), so treat ungrounded values as read from the image, not from cited text.
FUNSD here has no question-answer links, so this is entity F1, not pair F1.
4. First integration: HCC/RADV scanned charts
Same synthetic members as the HCC eval (synth.py), each note printed on its own page with a signature block (a pen
squiggle, or a blank line when unsigned), addenda with their own signature lines, then degraded like an office scan
(grey, rotation up to 1.2 degrees, speckle, blur, JPEG quality 55) into an image-only PDF (hcc/scan.py). The scanned
chart goes through POST /hcc/read-chart (document reader, then the same review); the text version through
POST /hcc/review, on the same server at the same time. Model: Qwen3.8-27B on the model's direct route, "samples" judgment
method, as in the published HCC eval (the model gateway was in a restart loop during these runs, see Caveats).
| dev (12 members, 48 codes) | test (50, 200) | fresh (50, 200), after fix 1 | fresh2 (50, 200), after fix 2 | |
|---|---|---|---|---|
| Verdict accuracy, text | 46 / 48 | 192 / 200 | 191 / 200 | 190 / 200 |
| Verdict accuracy, scanned | 46 / 48 | 187 / 200 | 190 / 200 | 189 / 200 |
| Same verdict, text vs scanned | 48 / 48 | 195 / 200 | 197 / 200 | 195 / 200 |
| False keeps on scans (unsupported or held called supported) | 0 | 0 | 1 (text: 0) | 0 |
| Header fields read right (date, record type, provider, credential, signed) | 44 / 44 each | 181 / 182 each | 187 / 187 each | 166 / 166 each |
| Note text edit distance (mean) | 0.0002 | 0.0121 | 0.0018 | 0.0011 |
| Addenda read exactly | 4 / 6 | 15 / 18 | 19 / 23 | 19 / 20 |
| Members whose record checks match the text version exactly | 12 / 12 | 49 / 50 | 49 / 50 | 49 / 50 |
| Quotes cited to page + box | 73 / 73 | 315 / 315 | 336 / 336 | 326 / 326 |
| Seconds per member (3 in parallel): scanned / text | 38 / 36 | 38 / 10 | 35 / 8 | 32 / 9 |
Each held-out set was run once, and each found one reading bug that the next set then measured after the fix:
- Test split → fix 1. Two of the five differences were one bug: on two unsigned pages Docling's reading order put the specialty value after the note body, and the body started after the last header value in reading order, so the note came back empty and its code became "no mention". Fix: the body starts below the header block, by geometry. The other three were rheumatoid arthritis "seropositive" wording (the HCC eval's known weak spot) and an audio-only call read as "less specific". All five deleted or held; none kept.
- Fresh split → fix 2. One false keep: a coder's addendum written five months after the visit ("Addendum after
chart review: morbid obesity ...") was read into the physician's note, so the obesity code was kept. The layout model
did not find the "Addendum dated ... by ..." heading as a region, and the addendum date was tied to the addendum's
signature line, which sits below the addendum text. Fix: the addendum starts at the earliest of its date's element, a
region beginning "Addendum", and the bottom of the note's own signature line; a page whose addendum text cannot be
separated is flagged (
addendum_not_isolated). Re-run on that member: all four codes right. The other two differences were a phone call's evidence the text path missed ("no mention", the scan got it right) and an unsigned note read as "less specific".
Fresh2, after both fixes: the five differences go both ways (two where the scan was right and the text path wrong); one is reader-related (a late addendum's text went missing, so its code was deleted rather than held: the safe direction).
A safety note that follows: on scans the no-false-keep property depends on the reader separating records (notes,
addenda) correctly, not only on the review. The per-page fields, body_elements and addendum_elements in the
/hcc/read-chart answer show how each page was split, so a coder can check it.
5. Production runtime: our server's GPU0 (vLLM), measured 27 Sep 2026
Sections 1 to 4 ran PaddleOCR-VL-1.6 on the Mac (mlx-vlm). Production runs it on our server: the same weights
(PaddlePaddle/PaddleOCR-VL-1.6@c5630ab, model.safetensors sha256 85a479d5…) in vLLM 0.29.0 (the pinned image the
language pack and the hosted model use), bf16, greedy, behind services/docreader with DOCREADER_PARSER=paddle-openai, and
Docling's layout model on CUDA. Two systemd user units, copies in deploy/systemd/: decosa-docreader-vlm (vLLM,
127.0.0.1:8498, --gpu-memory-utilization 0.04) and decosa-docreader (127.0.0.1:8497, the URL decosa-api already
defaults to). GPU0 is shared: about 74 GB of it was held by other services during every run below, so these are speeds
under load, not solo. Script: scripts/docreader_lab/beast_measure.sh, then compare_runtimes.py.
Accuracy: the same pages, the same Docling regions (the Mac's layout cache), the same scorers.
| Mac, mlx-vlm | Our server, vLLM | Regions read identically | |
|---|---|---|---|
| OmniDocBench dev, 30 pages: text NED / word F1 | 0.228 / 0.830 | 0.242 / 0.814 | 583 / 597 (97.7%) |
| OmniDocBench test, 84 pages: text NED / word F1 | 0.148 / 0.876 | 0.148 / 0.876 | 1,785 / 1,827 (97.7%) |
| OmniDocBench test, 24 tables: TEDS / TEDS-S | 0.818 / 0.858 | 0.833 / 0.874 | |
| FinTabNet dev, 40 tables: TEDS / numbers tied | 0.963 / 98.3% | 0.963 / 98.0% | 38 / 40 |
| FinTabNet test, 200 tables: TEDS / numbers tied | 0.930 / 83.8% | 0.932 / 84.4% | 182 / 200 |
| FinTabNet test: tables with every number tied | 147 / 196 | 147 / 196 | |
| IAM dev, 40 lines: CER (parser alone) | 4.9% | 5.1% | 38 / 40 |
| IAM test, 200 lines: CER (parser alone) | 4.7% | 4.6% | 194 / 200 |
They match: held-out numbers are within 0.002 NED, 0.6 points of numeric tie-out and 0.1 points of CER, in both
directions. The two runtimes do not produce the same tokens everywhere (91 to 98% of regions come out byte-identical):
bf16 kernels on Metal and CUDA round differently, and greedy decoding follows a different token wherever two are nearly
tied, which shows up as LaTeX delimiters (\(…\) vs $…$), a dropped check-mark glyph or one digit in a long table. The
one visible gap, dev text NED 0.228 vs 0.242, is a single page: a region of garbled pinyin on a Chinese exam scan that both
runtimes misread, where vLLM fell into a repetition loop (4,088 characters against 43). Without that page the dev scores
are equal. That region comes back with hit_limit: true (2,048 tokens), which makes the block re-read it with Qwen3.8
(reader.py: a region is kept as read only when its mean token probability is at least 0.90 and it did not hit the
limit), so the pipeline does not pass the loop through silently.
Our server's own layout (Docling on CUDA) against the Mac's layout cache (MPS), 30 dev pages: 636 of 639 regions have a Mac region of the same kind at IoU > 0.9, 28 of 30 pages have identical regions in identical order. End to end through the service with our server's own layout, OmniDocBench test scores text NED 0.141, word F1 0.884, TEDS 0.833 (dev: 0.229, 0.823, 0.552). The dev TEDS drop is one page, a dense capacitor spec table (about 3,000 tokens): the CUDA layout put its box one pixel to the left, and that crop changed the greedy reading of the table (TEDS 0.64 → 0.44). Long dense tables are where the reading is least stable, which section 2 already says for scans.
Speed (seconds per page, the service end to end: layout + reading, one page at a time):
| Our server (GPU0 shared, vLLM) | Mac | |
|---|---|---|
| The bench's 6 OmniDocBench dev pages, 50 regions: reading only | 1.56 (540 tokens/s) | 3.28 solo (258 tokens/s) |
| The same 6 pages through the service (layout + reading) | 1.70 | about 3.5 to 4.3 [layout not timed on these pages] |
| OmniDocBench dev, 30 pages, through the service | 2.62 (layout 0.38; slowest page 9.0) | 8.4 (reading only, shared GPU) |
| OmniDocBench test, 84 pages, through the service | 2.41 (layout 0.28; slowest page 8.5) | 6.0 (reading only, shared GPU) |
| The rehearsal fixtures (4 scanned chart pages + the born-digital page) | 0.86 | |
Layout only (read: "none"), 30 dev pages |
0.38 | 0.15 to 1 |
The first page after a restart pays about 6 s for loading the layout model. vLLM batches, but the service reads one
region at a time, so these are sequential numbers; decosa-api runs at most DECOSA_DOCREADER_MAX_CONCURRENT (2) documents
at once and each waits for the service's page lock.
GPU memory (nvidia-smi every 0.5 s through all of the runs above, 1,281 samples): vLLM 4,416 MiB peak (3,994 at
start: weights 1.82 GiB, activations 0.53, CUDA graphs 0.10, KV cache 1.16 GiB = 67,360 tokens, 8 requests of 8,192);
the service with the layout model 1,006 MiB; 5,422 MiB (5.7 GB) together, under the block's 6 GB ceiling on GPU0.
At --gpu-memory-utilization 0.05 vLLM held 4,916 MiB with a 2.04 GiB KV cache (119,040 tokens) that one-page-at-a-time
reading never fills; 0.04 keeps the total under the ceiling with room for the vLLM growth seen under load (+420 MiB).
End to end in decosa-api (27 Sep, the live API on :8445; no restart and no env change were needed, since
DECOSA_DOCREADER_URL defaults to http://127.0.0.1:8497): /docreader/info reports the service reachable; the
document-reader smoke passes (born-digital sample, 498 ms); POST /docreader/read {"sample": "scanned-chart", "forms": true} reads 4 pages from pixels (91 elements, 27 form fields, 10 Qwen re-reads) in 15.3 s and its receipt verifies;
POST /hcc/read-chart {"sample": "scanned-mixed-file", "review": true} reads and reviews the chart in 26.7 s (41 s with
the Mac serving the reader) with the same eight verdicts as the Mac-served run (3 not supported, 2 insufficient, 3
supported) and 17 of 17 quotes cited to page and box. The GPSR listing pack does not call the document reader yet: its
label image goes to Qwen3.8's vision input (section "Where it fails" does not cover GPSR).
Where it fails
- Tables on scans (section 2). Structure errors (merged or split columns) as much as digit errors.
- Charts: figures are only located unless
charts: true; then Kaplan-Meier, forest, bar and line charts are read into numbers by the chart reader, measured in docs/evals/chart-reader.md. - Two stacked tables inside one layout region are read as one (dev, OmniDocBench): a layout error the parser cannot fix.
- Handwriting: 4.0% CER with re-reads is fine for search and review, not for a number or a name without a person looking.
- Receipts: half of the values Qwen reads are not in the text of the element it cites.
- Scanned-chart eval pages are clean compared with real faxes (one font, no skew over 1.2 degrees, no handwriting in the body, no stamps or overlapping marks). Real charts will read worse.
Caveats
- English only. No language other than English was measured.
- The A/B used crop-level reading under one layout (Docling) for all four parsers; their own page-level pipelines (the numbers on their model cards) were not run.
- The IAM, CORD and FUNSD Qwen runs, and the HCC runs, used the Qwen3.8 model's direct route with this instance's own receipts, because the model gateway was being killed and restarted (SIGKILL, restart counter 4) from about 20:00 on 26 Sep while other jobs held most of our server's memory. Same model and weights; receipts are "attested", not gateway signed. The IAM dev Qwen run went through the gateway (signed).
- Speed on the Mac was measured with other jobs on the GPU except where marked solo. Speed on our server (section 5) was measured with about 74 GB of GPU0 held by other services and their load unknown; nothing was measured solo there.
- Everything in the HCC comparison is synthetic, written by the same agent that wrote the reader.
Rehearsal: checkable properties (rehearsal/document-reader/expected.json)
- The born-digital page is read from its text layer; its table comes back as one 7-row table.
- All 22 numbers the parser reads in that table are in the PDF's text layer.
- The scanned chart is 4 pages, all read from pixels, every element with page and box; at least 4 form fields tied to elements.
- The parser is PaddleOCR-VL-1.6 at a pinned revision; the document receipt is
decosa.docreader-receipt.v1and verifies. - Removing page 1's text elements makes the receipt fail.