Skip to content
decosa

The full write-up behind the numbers on the tool’s page, as the team that built it wrote it: data, method, results and caveats. Internal names are removed; nothing else is edited.

Document reader (block): eval

Date: 26 Sep 2026; production runtime on our server measured 27 Sep 2026 (section 5). Block: decosa_api.docreader + services/docreader. Lab code: scripts/docreader_lab/. Every number below is in docs/evals/document-reader/{ab-dev,test,bench,hcc-dev,hcc-test,hcc-fresh,hcc-fresh2,our server}.json.

What was built and what was measured

Pipeline: Docling (MIT) finds and orders the regions of a page (layout model docling-project/docling-layout-heron @8f39ad3, Apache-2.0); a ~1B page parser reads each region (text, or a table as HTML with spans); Qwen3.8-27B re-reads a region when the parser's mean token probability is under 0.90 and reads form fields tied to element ids. Born-digital PDF pages take their text from the PDF's own text layer; only tables are read from pixels, and every number in them is checked against that layer.

The parser was chosen by an A/B on dev slices, then only the winner ran on held-out slices, once. All four candidates got the same Docling regions and their own documented element-level prompts, greedy decoding, bf16, on the Mac Studio (M3 Ultra). No prompt or threshold was tuned on a test slice. The one pipeline bug found on dev (reading order, below) was fixed before any test run. The HCC test split and the first fresh split each found one reading bug; each was fixed afterwards and measured on the next untouched split (section 4).

Licences (checked 26 Sep 2026)

Candidate Revision Licence Note
PaddleOCR-VL-1.6 (chosen) PaddlePaddle/PaddleOCR-VL-1.6@c5630ab Apache-2.0, LICENSE file in the repo 0.9B
MinerU2.5-Pro-2605 opendatalab/MinerU2.5-Pro-2605-1.2B@bff20d4 Apache-2.0 (card) Trap: the older MinerU2.5-2509-1.2B card now says AGPL-3.0; the MinerU application repo is not Apache either (GitHub: NOASSERTION). mineru-vl-utils is Apache-2.0. Pin the Pro-2605 weights.
TeleOCR StarDoc-AI/TeleOCR@a614331 Apache-2.0 per the card metadata; no LICENSE file in the repo remote code written for transformers 4.x (needs 4.57)
OvisOCR2 ATH-MaaS/OvisOCR2@1fc9221 Apache-2.0, LICENSE file 0.8B
Docling + layout docling 2.130.0; docling-layout-heron@8f39ad3 MIT + Apache-2.0
Datalab Marker / Surya / Chandra — Out: the weights forbid use that competes with Datalab's API not tested

Datasets (evaluation only; nothing redistributed, no images in this repo):

Set Licence Used
OmniDocBench (1,651 pages) repo Apache-2.0; the HF dataset card has no licence field; pages come from many publishers English pages only: dev 30, test 84 (stratified by source, seed 48)
FinTabNet (OTSL copy, test shard 0) card says "other"; IBM released FinTabNet under CDLA-Permissive [U: not re-read today] dev 40 tables, test 200 (196 of them have numeric cells)
IAM handwriting lines (Teklia copy) the copy is tagged MIT; the IAM database's own terms are non-commercial research [U: the IAM site did not answer today] dev 40, test 200 lines
CORD-v2 receipts CC-BY-4.0 dev 25, test 75
FUNSD forms research use per its authors [U: the licence text did not render on its site today] dev 15, test 35

IAM and FUNSD are fine for measuring; keep them out of anything we ship or sell.

1. Parser A/B (dev slices)

OmniDocBench text NED ↓ OmniDocBench word F1 FinTabNet TEDS FinTabNet numbers tied, cell for cell IAM CER
PaddleOCR-VL-1.6 0.228 0.830 0.963 98.3% (1,015 / 1,033) 4.9%
MinerU2.5-Pro-2605 0.235 0.805 0.973 87.1% 4.0%
TeleOCR 0.207 0.850 0.895 82.2% 12.0%
OvisOCR2 0.236 0.845 0.795 61.4% 3.6%
Tesseract 5.3.4 (today) 0.333 0.703 no tables — 42.0%

Speed and memory, solo (one parser at a time, nothing else on the GPU, 6 pages / 50 regions, Mac Studio M3 Ultra):

seconds per page (reading only) tokens/s peak memory runtime
PaddleOCR-VL-1.6 3.3 258 2.9 GB (MLX peak) mlx-vlm 0.7.3
MinerU2.5-Pro-2605 11.6 71 5.7 GB (MPS) transformers 5.17
OvisOCR2 22.5 44 6.8 GB (MPS) transformers 5.17
TeleOCR 23.7 43 7.5 GB (MPS) transformers 4.57

Plus Docling's layout model: about 0.15 to 1 s per page and under 1 GB. The runtimes differ: PaddleOCR-VL under transformers on MPS ran at about 7 tokens/s (an op falls back), so it runs on MLX; on the 5 dev pages both runtimes read alike (text NED 0.092 vs 0.093).

Choice: PaddleOCR-VL-1.6. The table numbers decide it: numeric grounding will quote numbers read from tables, and Paddle tied 98.3% of numeric cells on dev against 87% for the next. Its page text is within 0.02 NED of the best, it is 3.5 to 7 times faster on the Mac at half the memory, and its handwriting gap is what the Qwen3.8 re-read covers. TeleOCR reads page text best but tied 82% of table numbers and has 12% CER on handwriting; MinerU has the best table structure (TEDS) but places numbers in the wrong cell more often.

Bug found on dev: the first layout step took Docling's assembled-element order, which interleaves columns. Taking the order of Docling's document items fixed it (page NED 0.38 → 0.21 to 0.24 for every parser). The parsers' readings did not change (same regions), so the dev runs were re-sequenced, not re-run (scripts/docreader_lab/reorder_runs.py).

2. Held-out (PaddleOCR-VL-1.6, run once)

Result Tesseract on the same slice
OmniDocBench, 84 English pages: text NED 0.148 0.239
OmniDocBench: word F1 0.876 0.763
OmniDocBench: 24 tables, TEDS / TEDS-S 0.818 / 0.858 none
FinTabNet, 200 tables: TEDS / TEDS-S 0.930 / 0.944 none
FinTabNet: numeric cells tied, cell for cell 83.8% (4,488 / 5,353) none
FinTabNet: numbers read correctly anywhere in the table 94.7% none
FinTabNet: tables with every number tied 147 / 196 none
IAM, 200 handwritten lines: CER (Paddle alone) 4.7% 52.9%
IAM: Qwen3.8-27B alone 3.6%
IAM: the block (Qwen re-reads when Paddle's mean token probability < 0.90) 4.0%, 81 of 200 lines re-read
Seconds per page on the Mac (shared GPU during the run) 6.0 2.3

Numbers from tables are not yet trustworthy on their own. Page 48 set a kill bar: 99% of table numbers must tie out on FinTabNet before numeric grounding trusts them. Held out, Paddle ties 83.8% (dev looked better at 98.3% on 40 tables; the test slice is 5 times larger). About a third of the misses are misread digits (94.7% are read correctly somewhere in the table) and the rest are cells in the wrong row or column. So:

  • born-digital PDFs: every table number is checked against the PDF's own text layer (numbers.in_text_layer); a number that is not there is flagged;
  • scans: table numbers carry the parser's confidence and must be confirmed by a person before numeric grounding uses them. The block does not claim tie-out on scans.

3. Forms (Qwen3.8-27B, the block's forms prompt)

The block's method sends the page image plus the parser's elements (id: text) and returns fields tied to element ids. The comparison is the same prompt with the image alone (no elements, so no boxes).

dev test
CORD receipts, 9 total fields, field F1: block (image + elements) 0.854 0.875
CORD: image only 0.877 0.895
FUNSD forms, key F1 / value F1: block 0.747 / 0.621 0.586 / 0.615
FUNSD: image only 0.766 / 0.615 0.501 / 0.489
Filled values found in the text of the element they cite ("grounded") CORD 156 / 325; FUNSD 323 / 410

On short receipts the element list costs about 2 points of F1 against reading the image alone; on dense forms it helps (value F1 0.615 vs 0.489). The block keeps the element list because it is what ties each value to a page and box. A value that is not in its element's text is returned with grounded: false: on CORD that is about half the values (receipt prices are often split across regions), so treat ungrounded values as read from the image, not from cited text. FUNSD here has no question-answer links, so this is entity F1, not pair F1.

4. First integration: HCC/RADV scanned charts

Same synthetic members as the HCC eval (synth.py), each note printed on its own page with a signature block (a pen squiggle, or a blank line when unsigned), addenda with their own signature lines, then degraded like an office scan (grey, rotation up to 1.2 degrees, speckle, blur, JPEG quality 55) into an image-only PDF (hcc/scan.py). The scanned chart goes through POST /hcc/read-chart (document reader, then the same review); the text version through POST /hcc/review, on the same server at the same time. Model: Qwen3.8-27B on the model's direct route, "samples" judgment method, as in the published HCC eval (the model gateway was in a restart loop during these runs, see Caveats).

dev (12 members, 48 codes) test (50, 200) fresh (50, 200), after fix 1 fresh2 (50, 200), after fix 2
Verdict accuracy, text 46 / 48 192 / 200 191 / 200 190 / 200
Verdict accuracy, scanned 46 / 48 187 / 200 190 / 200 189 / 200
Same verdict, text vs scanned 48 / 48 195 / 200 197 / 200 195 / 200
False keeps on scans (unsupported or held called supported) 0 0 1 (text: 0) 0
Header fields read right (date, record type, provider, credential, signed) 44 / 44 each 181 / 182 each 187 / 187 each 166 / 166 each
Note text edit distance (mean) 0.0002 0.0121 0.0018 0.0011
Addenda read exactly 4 / 6 15 / 18 19 / 23 19 / 20
Members whose record checks match the text version exactly 12 / 12 49 / 50 49 / 50 49 / 50
Quotes cited to page + box 73 / 73 315 / 315 336 / 336 326 / 326
Seconds per member (3 in parallel): scanned / text 38 / 36 38 / 10 35 / 8 32 / 9

Each held-out set was run once, and each found one reading bug that the next set then measured after the fix:

  • Test split → fix 1. Two of the five differences were one bug: on two unsigned pages Docling's reading order put the specialty value after the note body, and the body started after the last header value in reading order, so the note came back empty and its code became "no mention". Fix: the body starts below the header block, by geometry. The other three were rheumatoid arthritis "seropositive" wording (the HCC eval's known weak spot) and an audio-only call read as "less specific". All five deleted or held; none kept.
  • Fresh split → fix 2. One false keep: a coder's addendum written five months after the visit ("Addendum after chart review: morbid obesity ...") was read into the physician's note, so the obesity code was kept. The layout model did not find the "Addendum dated ... by ..." heading as a region, and the addendum date was tied to the addendum's signature line, which sits below the addendum text. Fix: the addendum starts at the earliest of its date's element, a region beginning "Addendum", and the bottom of the note's own signature line; a page whose addendum text cannot be separated is flagged (addendum_not_isolated). Re-run on that member: all four codes right. The other two differences were a phone call's evidence the text path missed ("no mention", the scan got it right) and an unsigned note read as "less specific".

Fresh2, after both fixes: the five differences go both ways (two where the scan was right and the text path wrong); one is reader-related (a late addendum's text went missing, so its code was deleted rather than held: the safe direction).

A safety note that follows: on scans the no-false-keep property depends on the reader separating records (notes, addenda) correctly, not only on the review. The per-page fields, body_elements and addendum_elements in the /hcc/read-chart answer show how each page was split, so a coder can check it.

5. Production runtime: our server's GPU0 (vLLM), measured 27 Sep 2026

Sections 1 to 4 ran PaddleOCR-VL-1.6 on the Mac (mlx-vlm). Production runs it on our server: the same weights (PaddlePaddle/PaddleOCR-VL-1.6@c5630ab, model.safetensors sha256 85a479d5…) in vLLM 0.29.0 (the pinned image the language pack and the hosted model use), bf16, greedy, behind services/docreader with DOCREADER_PARSER=paddle-openai, and Docling's layout model on CUDA. Two systemd user units, copies in deploy/systemd/: decosa-docreader-vlm (vLLM, 127.0.0.1:8498, --gpu-memory-utilization 0.04) and decosa-docreader (127.0.0.1:8497, the URL decosa-api already defaults to). GPU0 is shared: about 74 GB of it was held by other services during every run below, so these are speeds under load, not solo. Script: scripts/docreader_lab/beast_measure.sh, then compare_runtimes.py.

Accuracy: the same pages, the same Docling regions (the Mac's layout cache), the same scorers.

Mac, mlx-vlm Our server, vLLM Regions read identically
OmniDocBench dev, 30 pages: text NED / word F1 0.228 / 0.830 0.242 / 0.814 583 / 597 (97.7%)
OmniDocBench test, 84 pages: text NED / word F1 0.148 / 0.876 0.148 / 0.876 1,785 / 1,827 (97.7%)
OmniDocBench test, 24 tables: TEDS / TEDS-S 0.818 / 0.858 0.833 / 0.874
FinTabNet dev, 40 tables: TEDS / numbers tied 0.963 / 98.3% 0.963 / 98.0% 38 / 40
FinTabNet test, 200 tables: TEDS / numbers tied 0.930 / 83.8% 0.932 / 84.4% 182 / 200
FinTabNet test: tables with every number tied 147 / 196 147 / 196
IAM dev, 40 lines: CER (parser alone) 4.9% 5.1% 38 / 40
IAM test, 200 lines: CER (parser alone) 4.7% 4.6% 194 / 200

They match: held-out numbers are within 0.002 NED, 0.6 points of numeric tie-out and 0.1 points of CER, in both directions. The two runtimes do not produce the same tokens everywhere (91 to 98% of regions come out byte-identical): bf16 kernels on Metal and CUDA round differently, and greedy decoding follows a different token wherever two are nearly tied, which shows up as LaTeX delimiters (\(…\) vs $…$), a dropped check-mark glyph or one digit in a long table. The one visible gap, dev text NED 0.228 vs 0.242, is a single page: a region of garbled pinyin on a Chinese exam scan that both runtimes misread, where vLLM fell into a repetition loop (4,088 characters against 43). Without that page the dev scores are equal. That region comes back with hit_limit: true (2,048 tokens), which makes the block re-read it with Qwen3.8 (reader.py: a region is kept as read only when its mean token probability is at least 0.90 and it did not hit the limit), so the pipeline does not pass the loop through silently.

Our server's own layout (Docling on CUDA) against the Mac's layout cache (MPS), 30 dev pages: 636 of 639 regions have a Mac region of the same kind at IoU > 0.9, 28 of 30 pages have identical regions in identical order. End to end through the service with our server's own layout, OmniDocBench test scores text NED 0.141, word F1 0.884, TEDS 0.833 (dev: 0.229, 0.823, 0.552). The dev TEDS drop is one page, a dense capacitor spec table (about 3,000 tokens): the CUDA layout put its box one pixel to the left, and that crop changed the greedy reading of the table (TEDS 0.64 → 0.44). Long dense tables are where the reading is least stable, which section 2 already says for scans.

Speed (seconds per page, the service end to end: layout + reading, one page at a time):

Our server (GPU0 shared, vLLM) Mac
The bench's 6 OmniDocBench dev pages, 50 regions: reading only 1.56 (540 tokens/s) 3.28 solo (258 tokens/s)
The same 6 pages through the service (layout + reading) 1.70 about 3.5 to 4.3 [layout not timed on these pages]
OmniDocBench dev, 30 pages, through the service 2.62 (layout 0.38; slowest page 9.0) 8.4 (reading only, shared GPU)
OmniDocBench test, 84 pages, through the service 2.41 (layout 0.28; slowest page 8.5) 6.0 (reading only, shared GPU)
The rehearsal fixtures (4 scanned chart pages + the born-digital page) 0.86
Layout only (read: "none"), 30 dev pages 0.38 0.15 to 1

The first page after a restart pays about 6 s for loading the layout model. vLLM batches, but the service reads one region at a time, so these are sequential numbers; decosa-api runs at most DECOSA_DOCREADER_MAX_CONCURRENT (2) documents at once and each waits for the service's page lock.

GPU memory (nvidia-smi every 0.5 s through all of the runs above, 1,281 samples): vLLM 4,416 MiB peak (3,994 at start: weights 1.82 GiB, activations 0.53, CUDA graphs 0.10, KV cache 1.16 GiB = 67,360 tokens, 8 requests of 8,192); the service with the layout model 1,006 MiB; 5,422 MiB (5.7 GB) together, under the block's 6 GB ceiling on GPU0. At --gpu-memory-utilization 0.05 vLLM held 4,916 MiB with a 2.04 GiB KV cache (119,040 tokens) that one-page-at-a-time reading never fills; 0.04 keeps the total under the ceiling with room for the vLLM growth seen under load (+420 MiB).

End to end in decosa-api (27 Sep, the live API on :8445; no restart and no env change were needed, since DECOSA_DOCREADER_URL defaults to http://127.0.0.1:8497): /docreader/info reports the service reachable; the document-reader smoke passes (born-digital sample, 498 ms); POST /docreader/read {"sample": "scanned-chart", "forms": true} reads 4 pages from pixels (91 elements, 27 form fields, 10 Qwen re-reads) in 15.3 s and its receipt verifies; POST /hcc/read-chart {"sample": "scanned-mixed-file", "review": true} reads and reviews the chart in 26.7 s (41 s with the Mac serving the reader) with the same eight verdicts as the Mac-served run (3 not supported, 2 insufficient, 3 supported) and 17 of 17 quotes cited to page and box. The GPSR listing pack does not call the document reader yet: its label image goes to Qwen3.8's vision input (section "Where it fails" does not cover GPSR).

Where it fails

  • Tables on scans (section 2). Structure errors (merged or split columns) as much as digit errors.
  • Charts: figures are only located unless charts: true; then Kaplan-Meier, forest, bar and line charts are read into numbers by the chart reader, measured in docs/evals/chart-reader.md.
  • Two stacked tables inside one layout region are read as one (dev, OmniDocBench): a layout error the parser cannot fix.
  • Handwriting: 4.0% CER with re-reads is fine for search and review, not for a number or a name without a person looking.
  • Receipts: half of the values Qwen reads are not in the text of the element it cites.
  • Scanned-chart eval pages are clean compared with real faxes (one font, no skew over 1.2 degrees, no handwriting in the body, no stamps or overlapping marks). Real charts will read worse.

Caveats

  • English only. No language other than English was measured.
  • The A/B used crop-level reading under one layout (Docling) for all four parsers; their own page-level pipelines (the numbers on their model cards) were not run.
  • The IAM, CORD and FUNSD Qwen runs, and the HCC runs, used the Qwen3.8 model's direct route with this instance's own receipts, because the model gateway was being killed and restarted (SIGKILL, restart counter 4) from about 20:00 on 26 Sep while other jobs held most of our server's memory. Same model and weights; receipts are "attested", not gateway signed. The IAM dev Qwen run went through the gateway (signed).
  • Speed on the Mac was measured with other jobs on the GPU except where marked solo. Speed on our server (section 5) was measured with about 74 GB of GPU0 held by other services and their load unknown; nothing was measured solo there.
  • Everything in the HCC comparison is synthetic, written by the same agent that wrote the reader.

Rehearsal: checkable properties (rehearsal/document-reader/expected.json)

  1. The born-digital page is read from its text layer; its table comes back as one 7-row table.
  2. All 22 numbers the parser reads in that table are in the PDF's text layer.
  3. The scanned chart is 4 pages, all read from pixels, every element with page and box; at least 4 form fields tied to elements.
  4. The parser is PaddleOCR-VL-1.6 at a pinned revision; the document receipt is decosa.docreader-receipt.v1 and verifies.
  5. Removing page 1's text elements makes the receipt fail.