Skip to content
decosa

Developers · Building blocks

Document reader

Turns PDFs (born-digital or scanned), images, forms and handwriting into elements a checker can cite: text in reading order, tables as cells with row and column spans, form fields and checkboxes, each with its page and box, and a signed receipt over the page pixels and the model revisions.

Measured 2026-09-26. Full eval. First use: HCC evidence file (scanned charts).

Watch a real run

Two recorded runs: a synthetic chart printed, signed and scanned into an image-only PDF, and a born-digital PDF with a financial table. Every box is an element you can cite.

Loading the recorded run…

How it reads a page

  1. Docling (MIT) finds and orders the regions of each page with its Heron layout model (Apache-2.0).
  2. PaddleOCR-VL-1.6 (0.9B, Apache-2.0) reads each region: text, or a table as cells with spans. Chosen by an A/B of four ~1B parsers.
  3. Born-digital PDF pages take their text from the PDF's own text layer; only tables are read from pixels, and every number in them is checked against that layer.
  4. Where the parser is unsure (mean token probability under 0.90), Qwen3.8-27B re-reads the region from the image. Both readings are kept.
  5. Form fields: Qwen3.8 reads the page image with the element list and returns each value tied to an element; a value that is not in its element's text is marked as read from the image.
  6. decosa-api signs a decosa.docreader-receipt.v1: the file hash, each page's pixel hash, the layout and parser revisions, the hashes of the elements and fields, and the Qwen receipts.

Other blocks cite an element as p3 [120,488,940,66]: page, then the box in the page raster's pixels (PDFs at 144 DPI, so points are pixels divided by two). An element looks like this:

{
 "id": "p1-e16",
 "page": 1,
 "kind": "text",
 "bbox": [
  101,
  478,
  838,
  306
 ],
 "order": 15,
 "text": "Follow-up for diabetes, kidney disease and heart failure. Vitals: blood pressure 138/82 ...",
 "source": "parser",
 "conf_bp": 9958
}

API

POST /docreader/read{file_b64, synthetic: true, forms?, schema?, stream?} or {sample} -> pages of elements, fields, models, a signed receipt. Any valid session token or dk_ key.
POST /docreader/verify{document} -> signature, signer, elements, fields and page checks. No key.
GET /docreader/infoModels and pinned revisions, limits, whether the parser service is reachable, what it does not do.
GET /docreader/samples/{id}The synthetic sample PDFs: scanned-chart, born-digital.
POST /hcc/read-chartThe first integration: a scanned chart read into notes, then the HCC review with every quote cited to page and box.
POST /docreader/chart{image_b64, synthetic: true, method?: hybrid | model | geometry, hint?} or {sample} -> the chart as data (KM survival, medians, numbers at risk, HR; forest rows; bar and line values), plot box, method, confidence, receipts, a signed chart receipt.
POST /docreader/read {charts: true}Figure regions are read by the chart reader; each figure element gets its chart, covered by the document receipt.
POST /docreader/chart/verify{chart} -> signature, signer, table and image checks. No key.

Hosted, the reader takes synthetic or public documents only; self-host it for real ones. Limits: 12 MB and 10 pages per call on the hosted instance.

Which page parser, and why

Dev slices, same Docling regions for every parser, each parser's own element-level prompts, greedy, bf16, Mac Studio (M3 Ultra). Page text: 30 English OmniDocBench pages. Tables: 40 FinTabNet tables. Handwriting: 40 IAM lines.

ParserPage text NED ↓Word F1Table TEDSTable numbers tiedHandwriting CERs / page (solo)Peak memory
PaddleOCR-VL-1.6 (chosen)0.2280.8300.96398.3%4.9%3.3 (MLX)2.9 GB
MinerU2.5-Pro-26050.2350.8050.97387.1%4.0%11.65.7 GB
TeleOCR0.2070.8500.89582.2%12.0%23.77.5 GB
OvisOCR20.2360.8450.79561.4%3.6%22.56.8 GB
Tesseract 5.3.4 (what we used before)0.3330.703no tables—42.0%——

Numeric grounding quotes numbers read from tables, so tied table numbers decide it: PaddleOCR-VL tied 98.3% of numeric cells on dev against 87% for the next. Its page text is within 0.02 of the best, and it is 3.5 to 7 times faster on the Mac at half the memory. Its handwriting gap is what the Qwen3.8 re-read covers. Speeds are for different runtimes: PaddleOCR-VL on MLX, the others on transformers (PaddleOCR-VL under transformers on the Mac ran at about 7 tokens a second).

Held-out results

PaddleOCR-VL-1.6, run once after the dev work was frozen.

TestResultFor comparison
Page text, 84 English OmniDocBench pages: NED / word F10.148 / 0.876Tesseract: 0.239 / 0.763
Tables on those pages (24): TEDS0.818Tesseract: no tables
FinTabNet, 200 tables: TEDS0.930
FinTabNet: numeric cells tied, cell for cell83.8%94.7% of numbers read right somewhere in the table
IAM, 200 handwritten lines: CER4.0% with Qwen re-reads (81 lines re-read)PaddleOCR-VL alone 4.7%; Qwen3.8 alone 3.6%; Tesseract 52.9%
CORD, 75 receipts, 9 total fields: F10.875Qwen3.8 on the image alone: 0.895 (but no boxes)
FUNSD, 35 forms: key F1 / value F10.586 / 0.615Qwen3.8 on the image alone: 0.501 / 0.489

Table numbers read from scans are below the 99% bar we set before trusting them in numeric grounding (83.8% held out). Born-digital tables are checked number by number against the PDF's text layer; numbers from scanned tables go to a person.

First use: HCC charts, scanned vs text

Synthetic HCC members (invented people and notes), each note printed on its own page with a signature block, scanned into an image-only PDF, and reviewed; the text version of the same members through the existing path, on the same server at the same time.

MeasureResult
Verdict accuracy, scanned vs text (50 members, 200 codes, held out)187 / 200 vs 192 / 200
Same verdict, scanned vs text195 / 200
False keeps on scans (unsupported or held codes called supported)0
Header fields read right (date, record type, provider, credential, signed)181 / 182 each
Quotes cited to page and box315 / 315
Time per member (3 in parallel)about 38 s scanned, 10 s text
Later held-out splitsResult
Fresh split 1 (after fix 1): scanned vs text190 / 200 vs 191 / 200; 1 false keep on scans (a coder's addendum read into the note), fixed
Fresh split 2 (after fix 2): scanned vs text189 / 200 vs 190 / 200; same verdict 195 / 200
Fresh split 2: false keeps on scans0
Fresh split 2: header fields read right; quotes cited166 / 166 each; 326 / 326

Each held-out split was run once and each found one reading bug, fixed before the next: on the test split a note body dropped when the layout's reading order put a header value after it (codes deleted, never kept); on the first fresh split a coder's addendum whose heading the layout missed was read into the physician's note, and its code was kept. After both fixes, 50 new members: no false keeps. On scans, never keeping a code it should not depends on the reader separating notes and addenda correctly, so each page's split is shown in the answer.

Charts back into numbers

Figures back into numbers, so what a report says about a chart can be checked: Kaplan-Meier curves (survival at a time, medians, numbers at risk, a printed hazard ratio), forest plots (estimate and 95% CI per subgroup), and bar and line charts in annual reports and lay summaries.

  1. Qwen3.8 reads the labels and structure: tick labels, legend names and colours, row labels, printed numbers. One receipted call per chart.
  2. Code measures the geometry against those tick labels: the plot box and tick marks, each series by its colour, KM step curves traced left to right, forest markers and whiskers, bar edges to a fraction of a pixel.
  3. The page parser (PaddleOCR-VL-1.6) reads the number-at-risk table and printed forest values. Printed numbers are kept when the drawing agrees with them; when it does not, the row is flagged.
  4. If the axes cannot be calibrated, the model's own reading is returned and marked as such.
A synthetic Kaplan-Meier plot: two arms, 36 months, a number-at-risk table and a printed hazard ratio

The sample KM plot (fictional trial), read by the hosted API on 27 Sep 2026.

Read from the figureChart readerTrue
Median PFS, zenavotide14.48 months14.5
Median PFS, placebo8.46 months8.4
PFS at 12 months, zenavotide / placebo55.5% / 36.0%55.6% / 36.0%
Numbers at risk (14 cells)all exact
Hazard ratio (printed)0.62 (0.50 to 0.77)0.62 (0.50 to 0.77)

Held out: 240 synthetic charts (60 per kind; half clean, a quarter JPEG, a quarter scans), scored once. Same charts for every method.

MethodKM survival within 2 pointsKM median within 1% of the axisNumbers at risk exactForest: estimate and CI within 3%Bars within 2%Lines within 2%
Qwen3.8 alone58.7%13.1%97.1%62.4%73.6%81.9%
PaddleOCR-VL chart task59.0%23.4%48.7%58.7%44.2%76.4%
Code only (+ page parser)61.9%59.9%52.1%87.8%74.3%82.2%
Hybrid (chosen)84.5%79.6%89.6%90.8%88.5%92.1%

A log-axis fix found on the test set lifts hybrid forest rows to 99.8% (re-scored once after the fix, so not a clean held-out number). On clean charts hybrid medians are within 1% of the axis 88% of the time; on scans KM survival within 2 points drops to 74%. DePlot (Apache-2.0), on 60 of the same charts: 28% of KM survival values within 2 points, 0% of forest rows, 21% of bars; good only on simple lines (88%).

Five US public-domain figures (NCHS Data Briefs 480 and 492; an FDA statistical review, NDA 207103), labelled from the numbers the same documents print.

FiguresResult
NCHS bar charts, 3 figures, 40 barsThe model read the printed values exactly on two of three; the geometry fell back each time (no tick marks on the category axis, ticks drawn inwards). Group header rows shifted the third chart's categories.
FDA KM plots, 2 figuresThree of four printed medians found within 0.8% of the axis (the model alone: 3.6% to 14.7% off). A dashed red curve crossing a dashed black one was not traced.

In the CSR number-to-table verifier (figures: true), claims about KM medians, rates at a timepoint and hazard ratios, overall or for a subgroup's forest row, are checked against the figures within what the drawing can resolve. On the fictional ZEN-301 samples, the planted placebo median (9.7 stated, the curve crosses 50% at 8.5) and a subgroup HR (0.77 stated, 0.71 printed) were flagged and the 7 true claims were consistent, from JSON and from a PDF page where the document reader found the figure. A demo on our own samples, not a measured catch rate.

No open chart model with a clean licence reads KM or forest plots well. Our generator makes unlimited charts with exact truth, so a small chart model trained on it (for example PaddleOCR-VL-1.6 on its chart task) is a candidate for our own open model; proposed, not trained yet. Measured 2026-09-27. Chart reader eval.

Where it goes next

  • HCC / RADV evidence filebuilt

    Scanned charts to notes, the same review, quotes cited to page and box.

  • CSR number-to-table verifierbuilt

    Clinical study reports as PDFs or scans: the narrative's paragraphs and every in-text table and TLF read into cells, so each number can be traced to its cell and checked in code. With figures: true, what the text says about Kaplan-Meier medians, rates and hazard ratios is checked against the figures, read back into numbers.

  • GPSR listing packnext

    Declarations of conformity and test reports read with a schema for manufacturer, responsible person, product identifiers, standards and warnings; every listing field traced to a box.

  • CMMC evidence mapnext

    SSP and policy binders: born-digital pages cost almost nothing (text layer); each practice's evidence cited to page and box instead of a file name.

  • Denial appeal packetnext

    Faxed denial letters: claim number, reason, deadline read with cites; a deadline read from the image alone is confirmed by a person.

  • SNAP pre-QCnext

    Pay stubs, statements and bills: amounts read with cites and recomputed by numeric grounding; amounts from scans go to a person. Self-host only (PII).

Self-host

Two processes: the docreader service holds the models; decosa-api calls it for each page, adds the text layer, re-reads and fields, and signs. On a Mac:

python3.12 -m venv .venv && . .venv/bin/activate
pip install -r services/docreader/requirements-mac.txt
hf download PaddlePaddle/PaddleOCR-VL-1.6 --revision c5630abae1d940eafe0697512a0325494b02ab42 --local-dir models/PaddleOCR-VL-1.6
DOCREADER_MODELS=$PWD/models DOCREADER_PARSER=paddle-mlx python services/docreader/server.py
DECOSA_DOCREADER_URL=http://127.0.0.1:8497 DECOSA_DOCREADER_SYNTHETIC_ONLY=0 python -m decosa_api

On an NVIDIA GPU, the parser under vLLM and the service beside it:

hf download PaddlePaddle/PaddleOCR-VL-1.6 --revision c5630abae1d940eafe0697512a0325494b02ab42 --local-dir /models/PaddleOCR-VL-1.6
docker run --gpus device=0 --ipc=host -p 127.0.0.1:8498:8000 -v /models/PaddleOCR-VL-1.6:/model:ro vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1 /model --served-model-name PaddlePaddle/PaddleOCR-VL-1.6 --trust-remote-code --max-model-len 8192 --gpu-memory-utilization 0.04 --max-num-seqs 16 --no-enable-prefix-caching --mm-processor-cache-gb 0 --generation-config vllm
docker build -t decosa-docreader services/docreader
docker run --gpus all --network host -e DOCREADER_PARSER_URL=http://127.0.0.1:8498/v1 -e DOCREADER_PARSER_MODEL=PaddlePaddle/PaddleOCR-VL-1.6 decosa-docreader

Then rehearse on the bundled synthetic PDFs before any real document: python scripts/rehearse.py document-reader. Both paths are measured. GPU (how the hosted API runs it since 27 Sep 2026: vLLM 0.29.0 on an RTX PRO 6000 shared with other services; systemd units in deploy/systemd): 5.7 GB peak for parser and layout together, 1.6 s a page reading the bench pages, 2.4 s a page end to end on the 84 held-out pages, and the same accuracy as the Mac on the same pages (text NED 0.148 on both, table numbers tied 84.4% vs 83.8%, handwriting CER 4.6% vs 4.7%). Mac (mlx-vlm, the eval runtime): 3.3 s a page for reading plus up to 1 s of layout; about 3 GB. --gpu-memory-utilization is a share of the whole card: 0.04 is 3.8 GiB of a 96 GB card, so raise it on a smaller one.

Licences and receipts

PaddleOCR-VL-1.6Apache-2.0 (LICENSE file), PaddlePaddle/PaddleOCR-VL-1.6 @ c5630ab
Docling and its Heron layout modelMIT; Apache-2.0, docling-layout-heron @ 8f39ad3
Qwen3.8-27BApache-2.0, already served
pypdfium2 / PDFiumApache-2.0 or BSD-3-Clause / BSD-3-Clause
Chart readernumpy and Pillow (BSD-style, MIT-CMU); models as above. DePlot (Apache-2.0) measured, not used. UniChart and ChartInstruct (GPL-3.0), ChartGemma (Gemma terms), ChartLlama (research only) left out
Not usedMinerU2.5-2509 (now AGPL-3.0; the Pro-2605 weights are Apache-2.0); Datalab Marker, Surya and Chandra (the weights forbid competing use)

The parse runs outside our gateway today, so the document receipt is signed by decosa-api's own key: an attestation by its operator. The gateway design for model-call receipts (the service signs what it read; the gateway, which sees the bytes, countersigns) is in decosa-api docs/design/model-call-receipt.md, not deployed.

What it does not do

  • English only; no other language was measured.
  • Numbers in scanned tables are not trusted without a person (83.8% tied held out).
  • Handwriting: 4.0% character error rate on IAM lines with re-reads. Fine for search and review, not for a name or a number unchecked.
  • Formulas are read as text, not checked. Photos and figures other than Kaplan-Meier, forest, bar and line charts are located, not described.
  • A signature is noted as present or absent, never verified.
  • The hosted API takes synthetic documents only. The parser service runs next to it since 27 Sep 2026; the consoles do not offer uploads of scans yet.
  • Chart values carry the drawing's resolution: a median read from a curve is within about 1% of the time axis on clean charts, not exact. Scans, curves told apart only by dash, and category axes without tick marks are read less well.
All building blocks