Notes from your own jottings (jottings-note): eval
Date: 28 Sep 2026. Code: decosa_api/verticals/jottings; eval script scripts/jottings_eval.py; handwriting renderer
scripts/jottings_render.py. The test split ran once on the frozen commit 4fda26a; a fresh split measured the fixes the
test found (b0dc689); the fixes the fresh split and the cold-user test found (d38c41a to badfee7) were then run on the same
fresh sessions again ("fresh, final code" below: no longer held out, since those fixes came partly from reading its errors). Every run below used Qwen3.8-27B through the model gateway (receipted), 3 sessions in
parallel, on a gateway shared with other workloads; development runs used the direct route.
1. Data (all synthetic)
| Split | Sessions | Written by | Used for |
|---|---|---|---|
| dev | 14 (3 as photos) | a writer agent from a trap spec | all prompt and code changes |
| test | 48 (12 as photos) | a second writer agent, same spec, never looked at before the run | the frozen measurement |
| fresh | 24 (6 as photos) | a third writer agent, spec plus "vary where times and dates appear" | measuring the fixes the test found |
| demo backlog | 20 | the dev writer, one fictional therapist's two weeks | the hosted catch-up sample only |
The spec (eval/WRITER-SPEC.md in the build notes) asks for realistic post-session jottings (shorthand, arrows, typos,
quotes, ratings) across individual CBT, DBT, couples, family, play therapy, EMDR, ACT, MI, ERP, TF-CBT, grief, IPT and
groups, in person, video and phone, and plants traps, each at least 4 times in the test split:
T1 a third party's opinion ("mom thinks he has ADHD"), T2 a client's self-diagnosis, T3 the clinician's own "? r/o",
T4 a joke about death or violence, T5 a behaviour that must not become an emotion ("cried", "on phone a lot"), T6 missing
or partial times, T7 no plan, T8 no interventions, T9 exact numbers and scores, T10 a follow-up date and time, T11 a
client's report of a prescriber's medication change, T12 homework not done. Each session has gold labels for the
elements present and for whether risk is mentioned. Photos are rendered with OFL/Apache handwriting fonts (Caveat, Kalam,
Reenie Beanie, Indie Flower, Patrick Hand, Covered By Your Grace) on a ruled page, with perspective, uneven light, blur,
noise and JPEG compression: not real handwriting.
2. The blind check (unsupported sentences)
A blind Claude Code sub-agent (Opus 5.5) judged every sentence and header line of each note against the jottings and
goals only, with eval/JUDGE-INSTRUCTIONS.md: SUPPORTED, UNSUPPORTED (added facts, interpretations, moods, diagnoses,
risk levels, progress judgements, recommendations, or a dropped qualifier such as a joke, "no intent" or "?"), or
CONTRADICTED, plus a flag for an added clinical claim. On the test split the packets mixed our notes with notes from a
plain one-prompt baseline (the same model asked to "write a DAP progress note for insurance documentation ... use
professional clinical language and tie it to the treatment-plan goals"), shuffled and unlabelled; 8 judge agents, 12
cases each. Placeholders ("[not in your notes]") are not sent to the judge.
| Units judged | Not supported | Unsupported | Contradicted | Added clinical claim | Notes with any problem | |
|---|---|---|---|---|---|---|
| Tool, test (48 sessions, frozen) | 701 | 25 (3.6%) | 16 | 9 | 5 | 18 / 48 |
| Baseline, test (same sessions, same model) | 1,574 | 751 (47.7%) | 743 | 8 | 454 | 48 / 48 |
| Tool, fresh (24 sessions, after the test fixes) | 326 | 5 (1.5%) | 4 | 1 | 1 | 3 / 24 |
| Tool, fresh sessions again, final code (not held out) | 332 | 2 (0.6%) | 2 | 0 | 0 | 2 / 24 |
The target was 0 unsupported sentences. It was not met. What the judge found on the tool's notes:
- Test, from code (header lines), fixed after the run: a time range from another line taken as the session time ("cravings worst 5-6p" became 5 PM to 6 PM); a caregiver segment ("CG 15 min") as the total time; an intake date ("was 17 at intake 9/2") and an upcoming anniversary as the date of service; four "Risk (as written)" lines built from jokes ("my boss is going to kill me") that the judge counted as added risk claims. Fixes: the session line is chosen by what is on it, context words exclude other times and dates, jokes are no longer risk mentions (the qualifier guard keeps them jokes in the body).
- Test, from the drafted sentences (about 12): who said or did something ("dad: missed school 2 days" became "the client reported"; a client's quote given to the therapist), a family relation swapped ("wife's dad" became "husband's father"), shorthand misread ("sat" as sitting instead of Saturday, "ph 3-4" as a property of the target). Fixes: an attribution guard and a family-relation guard in code, and the drafting rules; shorthand like "sat" is not fixed.
- Fresh (5): a rating ("interference 6/10") taken as the date; "sat" read as sitting again; "? r/o opioid misuse" written as ruled out; "better off gone" without its "no plan, no intent"; "drove 1 exit" as "Highway 1". Fixes after this run: a rating on the 10th is not a date; "?" and rule-out lines are held back from the draft and listed for the clinician; "better off", "passive", "item 9" and similar words now carry the no-plan/no-intent check. On the same sessions with the final code the judge found 2: "drove 1 exit" as "to exit 1" again, and a homework item assigned to the child when the jotting did not say who.
The baseline shows why the checks exist: on the same jottings, a plain prompt to the same model wrote mood and affect ("Affect was congruent with mood"), engagement ("engaged, cooperative, appropriate eye contact"), diagnoses ("panic disorder"), progress judgements and next-session plans that the therapist never wrote, and added years to dates.
3. Completeness check (elements) and risk
Eight elements per session (date, start time, stop time, modality, interventions, goals addressed, response, plan); signature is scored separately from the details.
| Test (48, frozen) | Fresh (24, after test fixes) | Fresh again, final code (not held out) | |
|---|---|---|---|
| Elements right | 379 / 384 (98.7%) | 177 / 192 (92.2%) | 186 / 192 (96.9%) |
| Missing elements caught (missing recall) | 56 / 57 | 26 / 26 | 26 / 26 |
| Present elements called missing | 4 / 327 | 15 / 166 | 6 / 166 (4 goals, 2 interventions) |
| Risk mention right | 45 / 48 | 24 / 24 | 24 / 24 |
| Signature (name and credential given) right | 48 / 48 | 24 / 24 | 24 / 24 |
- Test misses: goals not linked (3; the model did not tag them and the goal gate needs a shared word), EMDR phase work not tagged as an intervention (1), and one date taken from an upcoming anniversary.
- Test risk misses: the gold for the test split counted jokes as risk mentions ("the probate lawyer is going to be the death of me"); the tool did not. After the blind check called joke-based risk lines added claims, the rule became: risk means suicide, self-harm, harm to others or a risk screen, including denials, not jokes. The fresh split's gold uses that rule: 24 / 24.
- Fresh misses: 10 of the 15 are start or stop times written on a line of their own ("2:15-3:05p") or as "8:00a start" after the test fixes made the time picker stricter; a date alone on its line; 4 goals. Fixed afterwards. The date and time detection is plain code, so it can be re-scored offline on the same gold: dev 42/42, test 144/144, fresh 72/72 after the fix (these splits are no longer held out for this part).
4. Reading handwriting
| Test (12 photos) | Fresh (6 photos) | |
|---|---|---|
| Lines read right (characters, after joining wrapped items) | 99 / 99 | 51 / 51 |
| Character error rate over the page | 0.02% | 0.0% |
| Lines flagged "check the reading" | 16, all read right | 6, all read right |
| Pages whose line split matched the jottings | 6 / 12 | 2 / 6 |
Qwen3.8-27B reads the page; PaddleOCR-VL-1.6 (the document reader's page parser) reads each line again from a crop found by ink density. The document reader's own layout model found only some of the handwritten lines on these pages (dev), so the second reader reads line crops instead. On dev one real misread was caught by the second reader ("G1" read as "61"); on test and fresh there were no misreads to catch, and every flag was a false alarm (mostly "0" vs "O", or punctuation). Line splits differ when a wrapped item is read as two lines or two items as one; the content is the same. Rendered fonts are easier than real handwriting: this is not a measure of real-world reading.
5. Time and cost
| p50 | p95 | Cost per note p50 (list price) | Model calls (mean) | |
|---|---|---|---|---|
| Test, gateway, 3 in parallel, shared | 42.4 s | 64.0 s | $0.0051 | 15.1 |
| Test, typed / photo | 40.6 s / 50.2 s | 62.7 s / 58.8 s | ||
| Fresh, gateway, 3 in parallel | 28.7 s | 64.9 s | $0.0051 | 14.7 |
| Fresh again, final code, gateway, 3 in parallel | 38.7 s | 49.0 s | $0.0051 | 14.7 |
| cbt-panic sample, gateway, 5 runs one at a time (pre-release server) | 16.4 s | 16.6 s | $0.0047 | 13 |
| Dev, direct route (self-host style) | 7.8 s | |||
| Catch-up, 20 sessions, gateway (recorded run) | 222 s for all 20 |
Cost is at the gateway list price ($0.30 / $1.50 per million input / output tokens) from the metered token counts. The one-prompt baseline cost about $0.001 and 14 s per note: the checks cost about five times more and take longer.
6. Cold user
A blind sub-agent playing a solo LCSW in Illinois, 20 notes behind, wrote her own 20-session backlog and used the self-hosted tool through a command-line client. First look: would use it to start notes, not sign them as is; about $25 a month (at most $40) once fixed; checking a draft took 2-7 minutes against 8-10 to write one (her estimate), so about 80 minutes instead of 3 hours for the backlog. Problems she found and what changed:
- total time taken from "5 min break" and "SUDS ... at 10 min" (only a session length counts now);
- her "? r/o" and "? tue or wed - check calendar" lines went into the note (now held back and listed);
- goal lines ignored unless a jotting said "(G1)" (a two-shared-word link now counts too);
- a jotting id "J5" read as the number 5, which stripped most of a note on one run (ids are stripped and ignored);
- batch output without sources or removed sentences (now included);
- a BIRP marked complete with an empty Response section (a present element now shows its jotting as written);
- "bf" became "ex-boyfriend", "aligns with the goal" (both now need the jottings);
- the setup: "I don't have a server in the back room". Not fixed; see limits. Second look, after those fixes: "I'd use it on my real backlog now and pay $25/month", $35 once it stops pasting raw shorthand into the note and gets who-did-what right ("the therapist used TIPP" when the client did). Fixed after the second look: jottings that no kept sentence carries are listed beside the note instead of pasted into it, a line ending in "?" outside a quote is held back, and the grounding judge gets a key to common shorthand (ind, LVM, HW, sat). Not fixed: who did what in some sentences, "stuck pt" (a CPT stuck point), a non-answer counted as a response, no year on dates (it never adds one).
7. Expected properties of the rehearsal run (rehearsal/jottings-note/)
- The cbt-panic jottings have every audited element:
report.missingis empty. - The risk line is the jotting as written:
Risk (as written in your notes): "denies SI/HI". - Times come from the jottings, with the minutes computed:
Total time: 53 minutes. - The mother's opinion never becomes a diagnosis: the note does not say "diagnosed".
- The couples session is missing its start and stop times and its plan, and says
[not in your notes]for them. - A request with a
recordingfield is refused with 400; the signed record verifies; every model call is receipted.
8. Self-host check (28 Sep, final code)
Fresh clone of the branch on our server, api image built from docker/api/Dockerfile, compose api service on the host
network with a named data volume, direct route to the running Qwen3.8-27B, the document reader's page parser and the
language pack's ASR; torn down after. Rehearsal 12 / 12 four times in a row (8.5-9.6 s); smoke module ok (6.2 s, 17
signed receipts, $0.0058 at list price); the photo sample read 9 lines (8 agreed by both readers, 1 flagged); 4 synthetic
dictations (a computer voice, 17-32 s) transcribed and drafted, spoken dates and times read ("October fourteenth, one to
one fifty p.m." gave 1 PM to 1:50 PM); a 190 s recording refused.
10. Risk and safety statements, and the therapy chain's unsupported sentences (29 Sep 2026)
Agent fix-therapy-accuracy-opus. Why: the therapy practice workflow's blind review found 29 of 1,283 draft units not
supported by the jottings, and one draft sentence kept a client's quote ("better off w/o me") with "no plan or intent"
but without the written "no SI". Code: risk.py, guard.py (meaning checks), compose.py (drafting rules 10-15 and the
meaning-check prompt), elements.py, note.py; tests tests/test_jottings_risk.py, dev variants
tests/data/jottings_risk_dev.json.
The hard rule, in code. Every risk or safety statement in the jottings is carried with its meaning unchanged, or the draft is blocked and the clinician is shown exactly which statement is missing or changed.
- Rules find the statements: suicidal and homicidal ideation incl. denials in any spelling ("denies SI/HI", "SI -", "no si/hi/sh"), passive SI and thoughts of death, self-harm and urges, safety plans, access to means, risk screens, possible mandated-report matters, substance-use risk. Each has a status (denied, present, history, a question, a step taken, pending, declined) and the qualifiers that change the meaning ("passive", "no plan", "no intent", "superficial", "didn't act", "urges", "free since"). Jokes and figures of speech are not statements. The model's risk tag is a second look that can only add a jotting.
- The header quotes every jotting holding a statement word for word.
- A drafted sentence that touches a risk topic must keep the statement's status and qualifiers (a quote about death keeps the "no SI" written beside it); otherwise one rewrite, then it is removed.
- After assembly the rule runs again on the note and on every text handed out (EHR paste, per-section copy, markdown; in the practice, its EHR paste and any late entry). Anything missing or altered blocks the draft: no copy text, the message names each statement. The practice also refuses a signed EHR copy that lost a statement unless the clinician confirms the change.
Also changed (the 29, by class): whose relative or group, an added relationship, the session's work made the client's report, a held meeting made "scheduled", session times tied to one activity, an added comparison with a goal (goal text no longer counts as evidence for progress words), a dropped hedge, shorthand ("sat" is Saturday), "response and progress: [not in your notes]" beside a goal-progress line, a goal listed as addressed when its work was "not reached"; plus a meaning check per kept sentence (one model call: who, when, qualifier, judgement, shorthand, added detail). On 1,657 judged body units from earlier runs (dev), the code checks catch 15 of the 42 the blind judge had called unsupported and send 9 of 1,615 supported units to a rewrite; the meaning check catches 17 of 42 and flags 4 of 400 supported.
Fresh held-out split (60 sessions written blind before any tuning, 92 planted risk statements in 39 sessions, 22 decoys; frozen code 4bd5d3f1; gateway; run once).
| Result | |
|---|---|
| Units not supported (blind Claude Code Opus judge, jottings eval instructions) | 7 / 805 (0.87%): 7 unsupported, 0 contradicted, 0 added clinical claims; 7 of 60 notes. Target 0.5% not met. |
| ... the 7 | 3 medication changes the client reported ("per pt") written as fact; 2 goals listed as addressed from shared words (SUDS, diary card); 1 who ("Eli yelled"); 1 when ("next week same time") |
| Planted statements quoted word for word in the risk line | 86 / 92 (rules 85, the model's second look 1); the other 6 were in body sentences |
| Blind risk-carry check (judge listed 108 statements) | 106 carried, 2 changed, 0 dropped; 0 decoys presented as risk (0 of 22 quoted as risk by the tool). Both changes were body sentences ("SP reviewed" -> "the plan was reviewed"; a quote kept without its "? SI, screen next session"); the risk line quoted both word for word. Target 0 changed not met. |
| Drafts blocked | 0 of 60 |
| Elements right; missing caught; present called missing | 474 / 480; 54 / 55; 5 / 425 |
| Time and cost per note (gateway, shared) | p50 37 s, p95 70 s; $0.0072 mean, 23 calls (was $0.0051, 15 calls) |
After the run (fixes from reading its errors; not held out): rules for the patterns it showed (a safety plan not done yet, a dated cut, medication held, given away, saved or taken back, keys on drinking nights, detox, food insecurity, a past attempt, "what's the point" beside "? SI"); a question stays a question; "SP" is the safety plan; a "per pt" medication report stays the client's report; goal links skip rating words and jottings that name their own goal. The rules alone now find all 92 statement lines (decoys still 0 of 22). Redrafting the 8 sessions with a judged problem: both risk changes and 5 of the 7 unsupported units gone; "Eli yelled" and "next week same time" remain.
Therapy chain drafts, the same practice as §1.1 of page 82 (run gw3, gateway, 70 notes, blind judge): 13 of 1,292 units not supported (29 of 1,283 before), 1 contradicted (a goal listed as addressed beside "didn't get to sleep stuff"), 0 added clinical claims (2 before), 12 of 70 notes (23 before). Risk statements: 68 of 68 carried, 0 changed, 0 false alarms (before: 1 "no SI" dropped). Of the 13: 4 "Type of therapy and format: [not in your notes]" when a jotting named EMDR or Gottman (fixed afterwards in code: a named therapy or format is the modality), 1 unlabelled "[not in your notes]" line (now names its section), and 8 single sentences (shorthand, a hedge, a response placeholder, the goal header).
Round 3 (29 Sep): risk wording verbatim in body sentences, in code. A kept sentence that carries a detected risk or
safety statement must hold the jotting's own words; a paraphrase, even a faithful one, is replaced by code with
As written in the session notes: "<the jotting, word for word>"., a sentence stating risk its cited jottings do not hold
is removed, and the final check blocks a note that still holds a non-verbatim risk sentence.
- Fresh blind split #3 (40 sessions, 110 planted statements in 34 sessions: 22 negations, 11 passive SI with qualifiers, 18 safety plans incl. "SP", 11 harm to others, 14 self-harm incl. history, 34 others; 15 decoys; written before the rule; frozen code c07f06de; gateway; run once). Blind Claude Code Opus risk-carry check: 113 of 117 statements carried, 4 changed, 0 dropped, 0 false alarms. Target 0 changed not met. Drafts blocked 0 of 40; sentences removed by the rule 0. Sentences replaced by the jotting word for word: 100 of 395 kept (all 100 kept the meaning by our rules: the price of "no paraphrase"; 24 had no risk word and only shared a word with a risk jotting). Rules found 106 of 110 statements (the risk line quoted 104 word for word, plus "SI - HI -" with its spaces collapsed). $0.0082 a note, 26 calls, p50 28 s.
- The 4 changed: three sentences quoted only the client's words from a risk jotting ("I'd be happier with Joan", "wanted to wring her neck") or reworded another part of it ("meant the job") without the "denies SI" / "no plan" / "? SI" beside it; one safety step with no risk word ("added Thurs check-in + Wed phone check") that the rules do not detect.
- Fixed after (e2dc847a, not held out): a kept sentence that cites a jotting holding a risk statement must quote that jotting whole, word for word (on replay the three sentences are replaced; 12 more sentences of these drafts would be); the model's risk tag also covers a step taken because of risk; rules for the 4 lines the frozen rules missed and a dash read wrongly ("SI - HI -" was read as SI present). A fourth fresh split is needed to measure 0 changed on this code.
9. Limits
- Synthetic sessions written by agents, and rendered handwriting; no licensed therapist has read the outputs yet, and no real jottings or real handwriting were used. A licensed therapist should read 20 outputs before any claim beyond this.
- The judge is one model family (Claude Opus) reading synthetic notes; its strictness shapes the numbers (for example it calls a goal link "inferred" when the jotting does not name the goal).
- Not zero: 1.5% of units on the fresh split were still not supported, mostly shorthand read the wrong way ("sat"); 0.87% on the second fresh split (29 Sep, §10).
- Runs vary: the same jottings can give different drafts on two runs (temperature 0 on a shared, batched server).
- Dictation: 4 synthetic dictations (a computer voice) were transcribed and drafted on the self-host box; spoken dates and times are converted to digits; not measured on real voices.
- A solo practice needs a machine with a 32-96 GB GPU (or a confidential hosted tier with a BAA, not available yet).