Skip to content
decosa

The full write-up behind the numbers on the tool’s page, as the team that built it wrote it: data, method, results and caveats. Internal names are removed; nothing else is edited.

Eval: medical chronology with page cites (use case 73)

Run 27 Sep 2026 on our server. Runner: scripts/chronology_eval.py. Results: docs/evals/medical-chronology/{dev,test,test-novision}.json (scores, misses, extra lines) and *-outputs.json (every entry with its cites).

What was measured

Nine synthetic record packets for fictional patients (decosa_api/verticals/chronology/synth.py), 16 pages each from five providers: typed PDFs with a real text layer (PCP, specialist), scanned pages (ED note, lab table, operative note, one MRI report; rotation, speckle, blur, JPEG), a handwritten PT evaluation and flow sheet (OFL handwriting fonts with jitter), and faxed copies (heavier noise, a fax header with its own date). Each packet plants:

  • about 33 required events (encounters, diagnoses, procedures, imaging, labs, medications), each tied to the line(s) that state it and their boxes in page pixels;
  • 2 copied pages (the ED note faxed into the PCP's file, an MRI report faxed into the specialist's), and a PCP note that mentions the ED visit;
  • 1 date conflict (a post-procedure note gives the procedure a date 7 or 14 days off);
  • 1 gap in treatment of 95 to 140 days after the injury (gaps are recomputed from the planted care dates at 60 days);
  • 2 pre-existing diagnoses 1 to 3 years before the injury, one in a claimed body region (related), one not;
  • traps that are not events: negative findings, family history, denials, fax dates, signature dates, options discussed but not done. Referrals and orders, an operation note read as a visit, and an operation's indication are optional (neither credited nor penalised).

Splits: dev seeds 1-3 (used to change prompts, thresholds and the checks: five runs); test seeds 101-106, run once after the dev work was frozen. Before the test run, and after the last dev run, one answer-key change was made: the operative note's visit became an optional event (it was the only extra line on dev); no code or prompt changed. The blank-page handling (an empty report instead of a 400) was added after the test run and does not touch scored paths.

Scoring (code, no model): an entry matches a planted event of the same type when a key term of the event is in its label or quotes (best score: same date, more terms, a cite on a page that states it); each planted event matches at most one entry. Date right = same ISO date. Cite right = a cite on a page that states the event whose box overlaps the planted line's box by at least half the line height. An entry that matches an event another entry already took is a duplicate; one that matches nothing (not even an optional event) is "not a planted event".

Everything ran in process through the same Run the API uses, against the live document reader (:8497), the retrieval service (:8499) and Qwen3.8-27B through the model gateway (receipted), 4 page calls at a time, one packet at a time, while the gateway and GPUs were shared with other workloads.

Results

dev (3 packets, 48 pages) test (6 packets, 96 pages) test, no page image
Planted events found 98 / 99 189 / 198 (95.5%) 188 / 198 (94.9%)
Date right, of found 97 / 98 189 / 189 180 / 188 (95.7%)
Cite on the right page and box, of found 98 / 98 189 / 189 186 / 188
Lines that are not planted events 1 / 114 2 / 224 (0.9%) 9 / 230 (3.9%)
Duplicate lines (a merge that should have happened) 0 3 4
Events on faxed copies cited on the copy too 29 / 33 57 / 66 56 / 66
Date conflicts flagged (false) 3 / 3 (0) 6 / 6 (0) 6 / 6 (0)
Gaps in treatment (false) 3 / 3 (0) 6 / 6 (0) 5 / 6 (1)
Pre-existing flagged; related said right (false) 6 / 6; 6 / 6 (0) 12 / 12; 12 / 12 (0) 12 / 12; 12 / 12 (0)
Seconds per 100 pages (end to end) 461.0 537.7 758.6
Model calls / generated tokens 75 / 11,655 153 / 23,548 155 / 24,729

Time per 100 pages: 537.7 s (about 5.4 s a page) on the shared hosted gateway; packets took 64 to 141 s. The console's 16-page sample took 144 s streamed (reading 51.6 s, extraction 43.0 s, merging, judgments and the reranked conflict search 49.8 s) and 59-66 s when the reader's output for the sample was cached.

Where it fails (test)

  • Reader misses. 9 events missed: ED diagnoses and an imaging line on scans where the layout model found only the label ("Imaging:", "Diagnoses:") and not the text after it, and one MRI study; and 4 visits where the model listed the diagnosis and medication but not the visit itself. None of the misses was invented or wrongly dated.
  • Therapy techniques as procedures. Both lines that were not planted events are "manual therapy" and "therapeutic exercise" read from a PT plan. Defensible to a reviewer, but not what we planted.
  • Referral duplicates. The 3 duplicates are "Referred by Dr X" lines on the consult, a second referral line next to the PCP's own referral.

What the page image buys

Without the page image (text elements only) recall is about the same, but dates go wrong on 4.3% of found events (mostly the handwritten PT evaluation dated from another page, and lab and PT lines left undated), a handwritten flow sheet was lost in one packet (6 visits), one false gap appeared, and 9 lines were not planted events. The image costs prompt tokens (image tokens are billed) but not generated tokens. The hosted route sends it; "vision": false turns it off.

Checkable properties of the sample run (rehearsal bundle rehearsal/medical-chronology/)

  1. The rotator cuff repair (2024-04-28) is listed as a procedure and flagged conflict with the other date 2024-05-05.
  2. One gap in treatment: 2024-05-12 to 2024-09-27 (138 days).
  3. The 2022-11-15 low back pain diagnosis is flagged before the injury and judged related; hyperlipidemia unrelated.
  4. F1 p4 (the faxed ED note) and F5 p2 (the faxed lumbar MRI) are listed as copies of F2 p1 and F3 p1.
  5. Every entry has a cite with a box, every cite's quote is located in an index chunk, each file has a signed document-reader receipt, and the signed record verifies and fails when its entry count is changed.

Caveats

  • The same author wrote the generator, the answer key, the prompts and the scorer; one template family, tidier than real records (no stamps, no overlapping marks, one page per visit, no multi-visit narrative).
  • No real records and no nurse reviewer's chronology to compare with. Page 27's buyer claim (1,600 chronologies a week) is about volume, not about this tool.
  • The matcher is lenient on wording (key terms) and strict on type, date and box.
  • Timing is under shared load and varies 2x between runs; a private GPU runs faster.

Verdict (revised after the eval)

Would a buyer pay? For self-hosted first-pass chronologies with a cite on every line, plausibly yes: on these packets no found event was wrongly dated or mis-cited with the image on, only 2 of 224 lines were not planted events, and the conflicts, gaps and pre-existing flags came out right. The honest gaps are real-record messiness (stamps, multi-visit notes, 300-page hospital stays), reader misses on scans (the main cause of missed events), and throughput: about 5 s a page on a shared GPU means an hour for a 700-page packet. Missing before a paid pilot: a real, de-identified packet with a reviewer's chronology, a reader fallback for regions the layout model leaves as a bare label, billing records, and a reviewer accept/reject step per line.