103 · Healthcare · live
Notes from your own jottings
Eval results
Scored on a held-out or test splitRun 29 Sep 2026Eval write-up (decosa-api, access required)
- Risk and safety statements carried, fresh blind split #3 (verbatim rule)113 / 117held outn = 11740 sessions written blind, 110 planted statements, 15 decoys. 0 dropped, 4 changed (fixed after, not re-measured), 0 false alarms, 0 drafts blocked; 100 of 395 kept sentences replaced by the jotting word for word.
- Draft units not supported, fresh held-out split (29 Sep, blind review)7 / 805held outn = 8050.87%: 0 contradicted, 0 added clinical claims, 7 of 60 notes (3 medication changes the client reported written as fact). The 0.5% target was not met.
- Therapy practice drafts not supported (blind review, 70 notes)13 / 1,292syntheticn = 1,29229 / 1,283 before these fixes; 68 of 68 risk statements carried (before: one sentence dropped a written 'no SI').
- Cost per note at list price (mean), held-out split0.0072 USDheld outn = 6023 model calls: each kept sentence now gets a grounding check and a meaning check (was $0.0051, 15 calls).
- Draft units not supported by the jottings (blind review)25 / 701test splitn = 70116 unsupported, 9 contradicted; 5 flagged as an added clinical claim. 18 of 48 notes had at least one.
- Same, one plain prompt to the same model (baseline)751 / 1574test splitn = 1,574454 flagged as an added clinical claim; 48 of 48 notes had at least one.
- Draft units not supported, after the fixes (blind review)5 / 326held outn = 32624 new sessions written after the test run; 3 of 24 notes had at least one.
- Audited elements marked right379 / 384test splitn = 384Eight elements per session: date, start, stop, modality, interventions, goals, response, plan.
- Missing elements caught56 / 57test splitn = 57
- Handwritten lines read right99 / 99test splitn = 9912 synthetic photos rendered with handwriting fonts; not real handwriting.
Dataset
Synthetic post-session jottings written by writer agents from a trap spec: 14 dev, 48 test (12 as rendered handwriting photos), 24 fresh sessions written after the test run (6 photos); gold element labels by the writers. On 29 Sep a second fresh split of 60 sessions (92 planted risk statements, 22 decoys) was written blind before the risk rule was tuned.
Caveats
- Synthetic sessions written by agents from a spec the builder wrote; rendered handwriting fonts, not real handwriting.
- The unsupported-sentence numbers come from one blind reviewer model (Claude Opus); no licensed therapist has read the outputs.
- The fresh split measured the fixes the test found; later fixes (from the fresh split and a cold-user test) are checked only on dev and unit tests.
- Test risk-mention gold counted jokes as risk; the tool no longer does (the blind review called joke-based risk lines added claims).
- Latency was measured on a shared gateway under load from other workloads.
- The 29 Sep fixes that came from reading the second fresh split's errors are checked only on dev data and a redraft of those sessions.
- The verbatim rule's last fix (quote the whole risk jotting) came from reading split #3's errors; a fourth blind split is needed to measure it.
Nightly smoke check
Loading the nightly status…
- Result
- pass
- Run
- 30 Sep 2026
- Latency, this run
- n/a
- p50 over passed runs
- 25 s
- Receipts
- 24
- Model calls
- n/a
- Tokens
- n/a
- Cost per run
- $0.009
Self-host verification
Verified on 28 Sep 2026: fresh clone of decosa-api on our server, api image built from docker/api/Dockerfile, compose api with a named data volume on the host network, direct route to the running Qwen3.8-27B, page parser and ASR; local signing; torn down after
Rehearsal bundle 12/12 four times in a row (8.5-9.6 s); smoke ok in 6.2 s with 17 signed receipts; the photo sample read 9 lines (8 agreed by both readers); 4 synthetic dictations (computer voice, 17-32 s) transcribed and drafted, a 190 s recording refused; /jottings/app drafted a note at 1280 and 390 px with no horizontal scroll; a cold-user test ran a 20-session catch-up against it. Model-server startup was not re-run.
Rehearsal bundle: jottings-note.zip (3 KB, 12 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.
Known limits
- Hosted verification on production (decosa.ai, gateway route), 28 Sep 2026: the smoke (cbt-panic sample) run 5 times one at a time.
- Measured on synthetic sessions written by agents and on rendered handwriting; no real jottings, real handwriting or licensed therapist's review yet.
- Not zero: on 24 new sessions 5 of 326 draft units were still not supported by the jottings (mostly shorthand read the wrong way, such as 'sat' for Saturday).
- The same jottings can give different drafts on two runs.
- Dictation was tested with a computer voice only; spoken dates and times are converted to digits, other numbers stay as words.
- Needs a 32-96 GB GPU in the practice for real notes until a confidential hosted tier with a BAA exists.
Models and licences
Standard tier. Licence posture: permissive (Apache, MIT or BSD).
- Reading typed jottings, the completeness check (dates, times and risk in code), the number, clinical-claim, qualifier and attribution guards, layouts, catch-up and the signed record (no model; CPU)decosa-api jottings (decosa_api/verticals/jottings), importing the grounding judge (vertical 17)AGPL-3.0-or-later
- Drafts the sentences, tags the audited elements, checks every sentence (the grounding judge), rewrites a failed sentence once, and reads a photo's pageQwen3.8-27B (NVFP4)Apache-2.0
All quality evidence
Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.
Standard · one GPU for the model (hosted demo) (5)
- Draft sentences and header lines a blind reviewer found not supported by the jottings (48 held-out synthetic sessions): 25 / 701docs/evals/jottings-note.md, test split, 2026-09-28 (blind Claude Code Opus review)
- The same blind check on 24 new sessions after the fixes the test found: 5 / 326docs/evals/jottings-note.md, fresh split, 2026-09-28
- Same reviewer, the same sessions drafted by one plain prompt to the same model (no checks): 751 / 1574docs/evals/jottings-note.md, test split
- Audited elements marked present or missing correctly: 379 / 384docs/evals/jottings-note.md, test split
- Elements missing from the jottings that it marked missing: 56 / 57docs/evals/jottings-note.md, test split
Best · adds the second reader and dictation (3)
- Handwritten lines read right (12 rendered photos, test split): 99 / 99docs/evals/jottings-note.md, test split
- Lines flagged 'check the reading' that were in fact read right: 16 of 16 flagsdocs/evals/jottings-note.md, test split (the second reader's disagreements are mostly punctuation or letter case)
- Dictations of up to 60 s transcribed and drafted; a 190 s recording refused: 4 / 4; refuseddocs/evals/jottings-note.md, self-host check (synthetic voice)