Chart reader (document reader block): eval
27 Sep 2026. On a pre-release build. Code decosa_api/docreader/charts.py; lab scripts/chart_lab/; scores
docs/evals/chart-reader/*.json (every number below is in one of them).
The chart reader turns a figure back into numbers so they can be checked: Kaplan-Meier curves (survival at time t, medians, numbers at risk, a printed HR), forest plots (estimate and CI per row), and bar and line charts (a value per category). Five methods were measured on the same charts.
| Method | What reads the chart |
|---|---|
| model | Qwen3.8-27B through the model gateway, one call per chart, structured-output prompt (axis ranges, tick labels, series, values, the risk table, printed numbers). Receipted. |
| paddle | PaddleOCR-VL-1.6's own Chart Recognition: task (the document reader's page parser, already running; 0.9B, Apache-2.0). Returns a table. |
| deplot | google/deplot (Pix2Struct, 282M, Apache-2.0), on the Mac (MPS). Returns a table. |
| geometry | Code only for the geometry: plot box, tick marks, colour unmixing, KM step tracing by dynamic programming, forest markers and whiskers, bar edges to a fraction of a pixel. Tick labels, the risk table and printed forest numbers are read by the page parser (OCR: on crops). Series are known by colour only. |
| hybrid (chosen) | The model reads the labels and structure (tick labels, legend names and colours, row labels, printed numbers); code measures the geometry against the model's tick labels; the page parser reads the risk table and printed forest numbers. Printed numbers win when the drawing agrees with them; a disagreement is flagged. Where the geometry cannot be calibrated, the model's reading is returned and marked method: "model". |
Data
All synthetic sets are drawn by scripts/chart_lab/synth.py (matplotlib) with exact truth kept alongside: fictional
drugs, trials and companies, so the set is ours (AGPL-3.0-or-later with decosa-api). Styles vary per chart: 5 fonts, 4 sizes,
72 to 200 DPI, 5 palettes (including black and grey KM curves, and navy with light blue), line widths, grids, frames,
censoring marks (|, +, none), risk tables (coloured or black), legends, printed HR annotations, printed forest
values (half the forest plots), log and linear forest axes, value labels on bars, horizontal bars, rotated labels. Then
a quarter are saved as low-quality JPEG (q 35 to 70) and a quarter as scans (rotation up to 0.8 degrees, grey paper,
noise, specks, blur, JPEG q 40 to 65).
| Set | Charts | Seeds | Role |
|---|---|---|---|
| dev | 80 (20 per kind) | 1-20 | used to build and tune the geometry |
| test | 240 (60 per kind; 123 clean, 59 JPEG, 58 scans) | 100001-100060 | held out, scored once |
| test, DePlot subset | 60 (every 4th test chart) | DePlot is slow here (median 11 s per chart on MPS), so it ran on a subset | |
| real | 5 US public-domain figures | NCHS Data Briefs 480 (Figure 1) and 492 (Figures 1 and 2); FDA statistical review of NDA 207103 (Figures 5 and 6). Fetched 27 Sep 2026, cropped at 150 DPI. Labels are the numbers the same documents print (value labels, and the medians in Tables 9 and 10 beside the KM figures), checked by eye against the figures. |
On the test set: 1,190 survival-at-t points and 137 medians on 60 KM plots, 912 numbers-at-risk cells, 436 forest rows (256 with printed values), 541 bars, 1,087 line points.
Metrics: KM survival at t, absolute error in probability (every x tick inside follow-up plus two off-tick times); medians, error as a percentage of the time axis ("not reached" must match); numbers at risk, exact per cell. Forest: estimate and both CI bounds within 3% relative, and exact to 2 decimals where printed. Bars and lines: relative error.
Results
Held-out test (240 charts, scored once)
| KM: S(t) within 0.02 | KM: within 0.05 | KM: median within 1% of axis | median error, % of axis (median) | numbers at risk exact | Forest: all three within 3% | printed rows exact | unprinted rows within 3% | Bars within 2% | Bars within 5% | Lines within 2% | Lines within 5% | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| model (Qwen3.8) | 58.7% | 83.7% | 13.1% | 3.40 | 97.1% | 62.4% | 98.8% | 10.0% | 73.6% | 81.5% | 81.9% | 91.4% |
| paddle (chart task) | 59.0% | 86.1% | 23.4% | 2.58 | 48.7% | 58.7% | 95.7% | 5.0% | 44.2% | 48.8% | 76.4% | 85.1% |
| geometry (code + page parser) | 61.9% | 65.0% | 59.9% | 0.20 | 52.1% | 87.8% | 84.8% | 91.7% | 74.3% | 76.0% | 82.2% | 84.9% |
| hybrid | 84.5% | 87.7% | 79.6% | 0.21 | 89.6% | 90.8% | 98.8% | 78.3% | 88.5% | 92.8% | 92.1% | 94.8% |
Median relative errors (bars / lines): model 0.32% / 0.43%, geometry 0.22% / 0.19%, hybrid 0.08% / 0.17%.
By degradation (hybrid, test): KM S(t) within 0.02: clean 90.4%, JPEG 83.8%, scans 74.4%. Bars within 2%: clean 96.1%, JPEG 95.4%, scans 64.0%. Lines: clean 96.7%, JPEG 94.8%, scans 75.8%. Forest all three within 3%: clean 86.9%, JPEG 99.1%, scans 89.5%. Scans are where it is weakest.
A fix found on the test set (so the next row is not a clean held-out number): 9 of the 23 test charts where hybrid fell back to the model were forest plots with a log axis that the model called "linear"; calibration now tries the other scale when the named one does not fit. Re-scored once after the fix: forest all three within 3% 90.8% -> 99.8%, unprinted rows 78.3% -> 100%; the other kinds unchanged (fallbacks 23 -> 14). Dev was unchanged by the fix.
DePlot on a test subset (60 charts, same charts for every method)
| KM S(t) within 0.02 | KM median within 1% | NAR exact | Forest all three within 3% | Bars within 2% | Lines within 2% | |
|---|---|---|---|---|---|---|
| model | 62.3% | 21.9% | 93.3% | 74.5% | 67.1% | 82.0% |
| paddle | 60.9% | 21.9% | 67.1% | 67.3% | 37.0% | 76.1% |
| deplot | 27.7% | 12.5% | 0% | 0% | 20.5% | 87.5% |
| hybrid | 71.6% | 68.8% | 87.1% | 86.4% | 87.0% | 87.2% |
DePlot is good on simple line charts and poor at everything this block needs most (KM, forest, grouped bars). It is not used.
Dev (80 charts; what the code was tuned on)
| KM S(t) within 0.02 | KM median within 1% | NAR exact | Forest all three | printed exact | Bars within 2% | Lines within 2% | |
|---|---|---|---|---|---|---|---|
| model | 59.4% | 20.9% | 94.0% | 71.2% | 94.6% | 87.0% | 87.1% |
| hybrid | 96.1% | 97.7% | 99.6% | 99.3% | 94.6% | 98.1% | 93.8% |
Dev to test drops by 3 to 13 points for hybrid: the geometry was tuned on 80 charts and meets new styles on test.
Real public-domain figures (5)
| Figure | Kind | model | hybrid |
|---|---|---|---|
| NCHS 480 Fig. 1: long COVID by sex, 3 bars per group in 3 shades of blue, value labels | bar | 6/6 exact | 6/6 exact (fell back to the model: no tick marks on the category axis) |
| NCHS 492 Fig. 1: life expectancy, horizontal, 2 groups with sub-headers | bar | 12/12 exact | 12/12 (fell back: the value axis's ticks point inwards) |
| NCHS 492 Fig. 2: death rates, 11 horizontal pairs with group header rows | bar | 6/22 within 2% (the header rows shift the categories) | same (fell back) |
| FDA NDA 207103 Fig. 5: four PFS curves, black and red, solid and dashed | KM | medians 3.6% of the axis off | P+L median 18.1 read as 18.10 (0.01% of the axis); the LO curve was not traced (dashed red crossing dashed black) |
| FDA NDA 207103 Fig. 6: OS, 2 curves | KM | 14.7% of the axis off | 0.78% of the axis off, both medians found |
What the real figures show that the synthetic set did not: category axes without tick marks, tick marks drawn inside the axis, group header rows in a category list, and four curves in two colours told apart only by dash. The geometry fails safe on the first three (it falls back to the model's reading of the printed labels, which were exact on two of three bar charts) and fails on the fourth. Five figures are not a measurement; they are a list of the next things to build.
Choice
Hybrid, by measurement. The model reads text well (numbers at risk 97%, printed forest numbers 99%) and geometry badly (medians within 1% of the axis 13%); the code measures geometry well (medians 0.2% of the axis) but cannot name series or read labels without help. Together: medians within 1% of the axis for 80% (90% on clean charts), KM survival within 2 points for 85%, forest estimates and CIs within 3% for 91% (99.8% after the log-axis fix), bars and lines within 2% for 89% and 92%.
Numbers at risk: the page parser reads the table and the rows go to curves by name; when that fails, the model's reading is used. On test that gave 89.6% exact against the model's own 97.1%: the model reads the table better than the parser on dense or tightly spaced tables. Switching the risk table to the model's reading (with the parser as a cross-check) is the obvious next change; it was seen on test, so it is not made here.
Time and cost
Measured under the shared gateway's load, test set, 4 in parallel: the model call takes 9.3 s per chart (median), the page parser 0.8 s per call, the geometry about 0.6 s of CPU per chart. About 970 prompt and 440 generated tokens per chart: roughly $0.001 per chart at the gateway list price ($0.30 / $1.50 per million). The docreader routes add the page-parser calls (1 to 20 per chart; each 0.1 to 0.8 s).
Integration
POST /docreader/chart: one image in, the table out with its provenance (plot box, method per value, calibration residuals, confidence, receipts) and a signeddecosa.chart-receipt.v1.POST /docreader/read {"charts": true}: figure regions the layout model finds are cropped and read; the element carrieschart, covered by the document receipt. On the figure-page sample (a born-digital CSR page) the figure is found and read: medians 14.45 and 8.50 months against 14.5 and 8.4 true, numbers at risk exact, printed HR read.- CSR verifier
figures: true: claims about KM medians, rates at a timepoint and HRs (overall, or a subgroup's row on a forest plot) are listed by one receipted Qwen call (it sees the sentences, never the figures' numbers) and compared in code within the drawing's resolution (median: the larger of 1% of the axis, 2 px, and the stated rounding; rate: 2 points plus rounding; printed HR: exact at its decimals). Numbers no table holds becometraced_to_figureorfigure_mismatch. On the two samples (fictional ZEN-301): both planted errors flagged (placebo median stated 9.7, the curve crosses 50% at 8.5; ECOG 1 HR stated 0.77, printed 0.71), all 7 true figure claims consistent; on the JSON sample the verifier's flags fell from 20 to 3 (the two planted errors, and a "65" in "younger than 65 years" that its context rules do not skip) because the figure-only numbers were traced to their figures. On the PDF sample the figure was located by the document reader and read from the page. This is a demo on our own samples, not a measured catch rate. - Receipts: Qwen calls are gateway receipts; page-parser calls are
decosa.model-call.v1receipts (modelpaddleocr-vl-1.6, kindocr); the geometry is code, recorded by the reader version in the chart receipt.
Licences (checked 27 Sep 2026)
| Licence | Used | |
|---|---|---|
| Qwen3.8-27B (the hosted model) | as recorded for the hosted model on the other stacks | model and hybrid |
PaddleOCR-VL-1.6 @c5630ab |
Apache-2.0 (LICENSE file) | page parser; its chart task measured |
| google/deplot (Pix2Struct) | Apache-2.0 (card) | measured only |
| google/matcha-* | Apache-2.0 (card) | not run (DePlot's sibling, chart QA rather than derendering) |
| OneChart | Apache-2.0 (LICENSE) | not run [its base, Vary, not checked] |
| UniChart, ChartInstruct | GPL-3.0 | out |
| ChartGemma | MIT on the card over Gemma terms (the Gemma prohibited-use policy travels with the weights) | out |
| ChartLlama | code MIT, the repo says research use only | out |
| TinyChart, ChartMoE, StructChart | [unverified today] | not run |
| WebPlotDigitizer (AGPL-3.0), IPDfromKM (GPL-2) | copyleft | not used |
| Datasets: ChartQA (GPL-3.0, images from Statista, Pew, OWID), CharXiv (charts stay the authors'), ArxivQA (CC-BY-SA plus OpenAI terms), DVQA (CC BY-NC) | not used | |
| PlotQA (CC-BY-4.0 data), FigureQA (MIT), ChartX (CC-BY-4.0) | permissive | not used; candidates for training |
Everything in the eval is our own synthetic set or US federal public-domain figures.
Where it fails
- Scans: KM survival within 2 points drops to 74%, bars to 64%. Speckle is taken for tick marks (the snap-fit helps), and thin lines lose their colour to JPEG chroma subsampling.
- Curves told apart only by line style (the FDA figure's dashed red and dashed black), and black-and-grey KM plots.
- Category axes without tick marks, ticks drawn inside the axis, and group header rows in category lists (all three from the real figures): the reader falls back to the model's reading.
- Dense risk tables whose numbers run together: the parser reads them as one number; the reader then reads cell by cell (slower) or falls back to the model.
- The KM trace reads the step curve; a median where the curve sits on 50% for a stretch is the first time it reaches 50%, which is the usual definition but not the only one.
- Confidence is a rule of thumb (calibration residuals, trace coverage, model-geometry agreement), not a calibrated probability.
- Only the four chart kinds. Pie charts, stacked bars, box plots, waterfall plots and spider plots are not read.
Own-model candidate
No open chart model with a clean licence reads KM plots or forest plots well enough (DePlot 28% of KM survival values within 2 points; PaddleOCR-VL's chart task 59%). The synthetic generator makes unlimited KM, forest, bar and line charts with exact truth, including the hard styles the eval exposed. A small model trained on it would replace the part the model does worst (geometry) and could also learn series-by-line-style. Proposal (not done here): fine-tune PaddleOCR-VL-1.6 (0.9B, Apache-2.0, already deployed) on its chart task with 50k synthetic charts rendered to our JSON schema, on the Mac or our server; keep the code geometry as the check. Released under Apache-2.0 once approved.
Checkable properties (chart samples, rehearsal)
chart-km: medians within 0.7 months of 14.5 and 8.4; numbers at risk exactly those inkm.json; HR 0.62 (0.50 to 0.77) read from the annotation.chart-forest: every row's estimate and CI equal toforest.json, source "printed", agreeing with the drawing.chart-bar: all ten bars within 60 (1% of the axis) ofbar.json.- Every chart result carries a
decosa.chart-receipt.v1that verifies and fails once the table is changed. figure-pageread withcharts: true: one figure element with a KM chart; the document receipt verifies.- CSR
zenavotide-301-figureswithfigures: true: exactly the placebo median and the ECOG 1 HR flagged.