Skip to content
decosa
LiveHostedSelf-hostSelf-host first for real data

Write a note from my jottings

A draft DAP, SOAP or BIRP note from your own jottings, every sentence linked to the jotting it came from, and a list of what an auditor would find missing.

Measured7 / 805Draft sentences not supported by the jottings (blind check)
On production25 smedian on production (2026-09-30); slower when the service is busy
List price~$0.79 per 100 notesmeasured, at list price

Built on: Signed record

Loading the tool…

Use it your way

Use it from your codeThe hosted API with your key, and prompts to paste into a coding agent
Self-host · your GPUs · recommended

Run it yourself, on request

  • The same open models and app, on 1x RTX PRO 6000 (96 GB) or 1x RTX 5090 (32 GB, estimate) for Qwen3.8-27B; the photo's second reader needs about 3 GB more; typed jottings need nothing else.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Hosted · by Decosa

Get an API key

  • Call the notes from your own jottings API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Synthetic, public or test data only: real patient data belongs on your own hardware.

Build with it

Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.

Base URL
https://api.decosa.ai
Auth
Authorization: Bearer $DECOSA_API_KEY (or a demo session token)
Tool id
jottings-note

Use the hosted API

# Decosa "Write a note from my jottings": use the hosted API (made-up sessions only)

You are wiring Decosa's jottings-to-note tool into this project. It takes a therapist's own post-session jottings (typed
lines, a photo of handwritten notes, or a dictation of up to 60 seconds made after the session) and returns a draft DAP,
SOAP or BIRP progress note where every sentence cites the jottings it came from, the completeness check (date, start and
stop times, modality, interventions, goals addressed, response, plan, name and credential; risk when mentioned), the
sentences it removed and why, text to paste into an EHR, and a record signed by the server, with a signed receipt for
every model call. Use only what is listed below. If you need something else, stop and ask me.

- Base URL: `https://api.decosa.ai`
- Health check: `GET https://api.decosa.ai/healthz`.
- **The hosted API takes made-up sessions only.** Therapy notes are protected health information. Every request with your
  own jottings must carry `"synthetic": true`, and text that looks like a real identifier (an email, a phone number, an
  SSN, "DOB" or "MRN", a street address) is refused with 400. Real notes belong on a self-hosted box (see the self-host
  prompt). Never send real client information here.
- It never takes a session recording: a request with a `recording` field is refused, and audio over 60 seconds is refused.
  Do not build a recording feature around it. The note is a draft the licensed clinician reviews, edits and signs; do not
  send it to a client or sign it automatically.

## Auth: API key (or a demo session)
1. Preferred: an API key (`dk_…`) from "Get an API key" on the tool page, in `DECOSA_API_KEY`, sent as
   `Authorization: Bearer $DECOSA_API_KEY`.
2. Without a key: `POST https://api.decosa.ai/demo/session` with `{"vertical": "jottings-note"}` returns `{"token", "expires_at", "budget"}`.
3. One run at a time per demo token (409 while one is going); 402 when the token budget is spent.

## Endpoints
- `POST /jottings/note` (token). Body:
  ```json
  {"jottings": "9/24 2:03-2:56p video, pt at home\nind CBT - G1\n3 PAs this wk (was 5)\ndenies SI/HI\nHW: thought record daily",
   "format": "dap", "goals": ["G1: Reduce panic attacks from 5 a week to 1 a week"],
   "details": {"client_label": "Client A (made up)", "clinician": "Dana Reyes", "credential": "LCSW"},
   "synthetic": true, "stream": false}
  ```
  Instead of `jottings`: `image_b64` (JPEG, PNG or WebP up to 8 MB) or `audio_b64` (up to 60 s). `format`: dap, soap, birp.
  `{"sample": "cbt-panic"}` runs a built-in sample (also `dbt-photo`, `couples-gaps`).
- JSON answer: `{ready: {jottings: [{id, text, source, read?, check?}]}, elements: {elements: [{id, label, status, from,
  value, note}], risk, missing}, note: {note: {header, sections: [{id, title, items: [{text, from, status}]}], signature,
  held_back, left_out}, plain, sections_text, markdown}, report: {verdict, missing, removed: [{text, why}], held_back,
  left_out, to_check, receipts, record, usage}}`. `note.plain` is the text to paste; `note.markdown` shows each sentence's
  jottings in brackets.
- SSE (`Accept: text/event-stream` or `"stream": true`): `progress`, `ready`, `receipt`, `elements`, `sentence` (one per
  drafted sentence: kept, fixed or removed), `note`, `report`, `budget`, `done`.
- `POST /jottings/batch` (token): `{"sessions": [{"jottings": "...", "label": "...", "format": "dap", "goals": [...]}],
  "synthetic": true}`, up to 30 typed sessions → per session note and report, and a `batch_report` with what is missing
  across all of them.
- `POST /jottings/read` (token): a photo → the jottings read, each with its read-confidence, before drafting.
- `POST /record/verify` (no token): the `record` → `{ok, checks, summary}`. `GET /jottings/info`, `GET /jottings/samples`.

## Errors
400 bad input (the message names the field or the limit, says the session is not marked made up, or refuses a recording),
401/403 token, 402 budget, 409 a run already going on this demo token, 413 body too large, 429 busy (`Retry-After`).

Run it yourself (containers)

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

# Decosa "Write a note from my jottings": run it yourself (containers)

You are setting up Decosa's jottings-to-note tool on this machine, so therapy notes never leave it. It turns a
therapist's own post-session jottings (typed, a photo of handwritten notes, or a dictation of up to 60 seconds) into a
draft DAP, SOAP or BIRP progress note where every sentence cites the jottings it came from, lists what an auditor would
find missing, and signs a record with no note text. It never records a session. Nothing is sent to Decosa's hosted API.
The note is a draft the licensed clinician reviews, edits and signs.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/jottings-note.zip (3 KB, 12 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py jottings-note` (the api image carries the same bundle under /app/rehearsal/jottings-note/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py jottings-note --bundle jottings-note.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "every audited element is in the CBT jottings, so nothing is missing", "the risk line carries the jotting as written", "the start and stop times are the ones written, with the minutes between them"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
   instructions (docs.docker.com/engine/install). Install the NVIDIA container toolkit and check
   `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
   `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
   Read it. Keep the `llm` service (Qwen3.8-27B on vLLM, prefix caching on) and the `api` service. For the `api` service
   set `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b`,
   `DECOSA_JOTTINGS_SYNTHETIC_ONLY=0` (so this box accepts real notes) and bind every port to 127.0.0.1. For handwriting
   photos, optionally add a vLLM serving `PaddlePaddle/PaddleOCR-VL-1.6` and set `DECOSA_DOCREADER_PARSER_URL` to it;
   without it, set `DECOSA_JOTTINGS_SECOND_READER=0`. Never set the gateway route on this box: it would send PHI off it.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check.
4. Check: `curl -fsS http://127.0.0.1:<PORT>/jottings/info` shows `synthetic_only: false` and route `direct`;
   `GET /attest/signing-key` shows this box's public key. Show me the key.
5. Smoke test: get a token with `POST /demo/session {"vertical":"jottings-note"}` and send `{"sample": "cbt-panic"}` to
   `POST /jottings/note`. Expect nothing in `report.missing`, `Total time: 53 minutes` and the line
   `Risk (as written in your notes): "denies SI/HI"` in `note.plain`, and the mother's opinion kept as her opinion. Send
   `{"sample": "couples-gaps"}`: `report.missing` includes start_time, stop_time and plan. Send the first answer's
   `report.record` to `POST /record/verify`: `ok` must be true. A request with a `"recording"` field must return 400.
6. Open `http://127.0.0.1:<PORT>/jottings/app` in a browser on this machine: type or photograph jottings, Draft, Copy.
7. Report back: the public key, the smoke results and how long a note took.

Off by default. Never on a box that holds PHI. If I ask for it later, follow the Provide page instead of improvising.
Run it on your own hardwareWhat it needs, and the prompt that sets it up

Run it on your own GPU

Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.

  • CPU only, 64 GB RAMDoesn't fit

    Qwen3.8-27B (NVFP4) needs a GPU.

  • GeForce RTX 4090standard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)

  • GeForce RTX 5090standard tierRuns

    The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

  • 2x GeForce RTX 5090best tierRuns

    The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size). The best tier fits too.

  • L40Sbest tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU. The best tier fits too.

  • H100 80 GB (SXM)best tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU. The best tier fits too.

  • RTX PRO 6000 Blackwell 96 GBbest tierRuns

    The standard tier fits (57.6 of 96 GB). The best tier fits too.

  • 2x RTX PRO 6000 Blackwell 96 GBbest tierRuns

    The standard tier fits (57.6 of 192 GB). The best tier fits too.

  • Apple M3 Ultra (Mac Studio), 96 GBstandard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon.

  • Apple M5 Max, 64 GBstandard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon.

Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

  1. 1

    Check the GPU, Docker and the NVIDIA Container Toolkit

    The driver must see the GPU, and Docker must be able to pass it into a container.

    nvidia-smi
    docker compose version
    docker run --rm --gpus all ubuntu nvidia-smi
  2. 2

    Fetch the compose file

    One file describes the API, the speech model and the language model as services.

    mkdir -p ~/decosa && cd ~/decosa
    curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml
  3. 3

    Pull and start

    The first start downloads pinned model weights, tens of gigabytes.

    docker compose pull
    docker compose up -d
  4. 4

    Check health

    Wait until the API reports ok with both models loaded. Then point your app at the local base URL.

    curl -fsS http://localhost:<PORT>/healthz
    # {"ok": true, "asr": true, "llm": true, ...}
    curl -fsS -X POST http://localhost:<PORT>/demo/session \
      -H 'Content-Type: application/json' -d '{"vertical":"jottings-note"}'

Set up with a coding agent, rehearse on mock data, then go private

  1. Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
  2. Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's expected.json. Every check must print PASS.
  3. Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
Rehearsal command
docker compose exec api python scripts/rehearse.py jottings-note

Download the mock-data bundle (3 KB, 12 checks)expected.json

Two made-up therapy sessions. The first has every element an auditor looks for, plus a trap: the client's mother thinks she has bipolar disorder, which must stay the mother's opinion. The second has no start or stop time and no plan, and a joke about violence that must stay a joke (and is not a risk mention). The tool must draft a DAP and a SOAP note where every sentence cites a jotting, list the missing elements as [not in your notes] instead of filling them in, carry the risk line as written, refuse a session recording, and sign a record that verifies.

What the rehearsal checks
  • every audited element is in the CBT jottings, so nothing is missing
  • the risk line carries the jotting as written
  • the start and stop times are the ones written, with the minutes between them
  • the mother's opinion never becomes a diagnosis
  • no sentence was kept without a jotting behind it
  • the couples session is missing its start and stop times and its plan, and the draft says so
  • missing elements are marked, not filled in
  • the joke keeps its 'no intent' (in the note, or listed beside it as written)
  • a joke is not a risk mention
  • a session recording is refused
  • the signed record verifies
  • every model call has a signed receipt

Licence: Synthetic: the sessions, clients and clinicians are invented (decosa_api/verticals/jottings/data/samples.json). Part of decosa-api, AGPL-3.0-or-later.

Prompt for your coding agent

# Decosa "Write a note from my jottings": run it yourself (containers)

You are setting up Decosa's jottings-to-note tool on this machine, so therapy notes never leave it. It turns a
therapist's own post-session jottings (typed, a photo of handwritten notes, or a dictation of up to 60 seconds) into a
draft DAP, SOAP or BIRP progress note where every sentence cites the jottings it came from, lists what an auditor would
find missing, and signs a record with no note text. It never records a session. Nothing is sent to Decosa's hosted API.
The note is a draft the licensed clinician reviews, edits and signs.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/jottings-note.zip (3 KB, 12 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py jottings-note` (the api image carries the same bundle under /app/rehearsal/jottings-note/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py jottings-note --bundle jottings-note.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "every audited element is in the CBT jottings, so nothing is missing", "the risk line carries the jotting as written", "the start and stop times are the ones written, with the minutes between them"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
   instructions (docs.docker.com/engine/install). Install the NVIDIA container toolkit and check
   `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file:
   `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
   Read it. Keep the `llm` service (Qwen3.8-27B on vLLM, prefix caching on) and the `api` service. For the `api` service
   set `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b`,
   `DECOSA_JOTTINGS_SYNTHETIC_ONLY=0` (so this box accepts real notes) and bind every port to 127.0.0.1. For handwriting
   photos, optionally add a vLLM serving `PaddlePaddle/PaddleOCR-VL-1.6` and set `DECOSA_DOCREADER_PARSER_URL` to it;
   without it, set `DECOSA_JOTTINGS_SECOND_READER=0`. Never set the gateway route on this box: it would send PHI off it.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check.
4. Check: `curl -fsS http://127.0.0.1:<PORT>/jottings/info` shows `synthetic_only: false` and route `direct`;
   `GET /attest/signing-key` shows this box's public key. Show me the key.
5. Smoke test: get a token with `POST /demo/session {"vertical":"jottings-note"}` and send `{"sample": "cbt-panic"}` to
   `POST /jottings/note`. Expect nothing in `report.missing`, `Total time: 53 minutes` and the line
   `Risk (as written in your notes): "denies SI/HI"` in `note.plain`, and the mother's opinion kept as her opinion. Send
   `{"sample": "couples-gaps"}`: `report.missing` includes start_time, stop_time and plan. Send the first answer's
   `report.record` to `POST /record/verify`: `ok` must be true. A request with a `"recording"` field must return 400.
6. Open `http://127.0.0.1:<PORT>/jottings/app` in a browser on this machine: type or photograph jottings, Draft, Copy.
7. Report back: the public key, the smoke results and how long a note took.

Off by default. Never on a box that holds PHI. If I ask for it later, follow the Provide page instead of improvising.

Help me customise for my hardware

Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.

Hardware

GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page

RunsNotes from your own jottings on GeForce RTX 5090: use the Standard · one GPU for the model (hosted demo) tier

The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

What this tool's stack says about this hardware:

  • 1x RTX 5090 32 GB (fits): Estimate: Qwen3.8-27B NVFP4 needs about 20 GB of weights plus KV cache; not run for this use case.

Standard · one GPU for the model (hosted demo): what changesuses estimates

  • Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
Memory per component
  • Reading typed jottings, the completeness check: decosa-api jottings (decosa_api/verticals/jottings), importing the grounding judge (vertical 17). CPU. Runs on CPU (vram_gb 0 in stack.json).
  • Drafts the sentences, tags the audited elemen...: Qwen3.8-27B (NVFP4). ~57.6 GB (at least ~28 GB), weights 21.4 GB (from stack.json). Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)

Expected speed

Not measured.

Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.

Setup prompt for this hardware

The self-host prompt for Notes from your own jottings, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.

# Set up Notes from your own jottings on my hardware

Fetch https://decosa.ai/prompts/jottings-note-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied.

## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=jottings-note)

Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4).
Quality tier: Standard · one GPU for the model (hosted demo) (standard). Fit check: runs with changes, about 28 GB of 32 GB used; some memory numbers are estimates, not measurements.

First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything.

Use these components (the setup below describes the standard tier; change it to match):
- Reading typed jottings, the completeness check: decosa-api jottings (decosa_api/verticals/jottings), importing the grounding judge (vertical 17), CPU
- Drafts the sentences, tags the audited elemen...: Qwen3.8-27B (NVFP4) (nvidia/Qwen3.8-27B-NVFP4), 57.6 GB. Change: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

GPU placement (set each service's device and its vLLM --gpu-memory-utilization to about the share shown):
- GPU 0: Qwen3.8-27B (NVFP4) ~28 GB (88%); about 4 GB left

During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed.

The stack's own component list and compose layout: https://decosa.ai/prompts/jottings-note-assemble.md

The proof

How we tested itEval results and end-to-end checks, hosted and self-hosted, with dates

Verified end to end

Hosted: verified 30 Sep 2026 · measured 30 Sep 2026: · p50 25 s · p95 30 s (5 runs) · ~$0.009 per run · 24 receipts

Loading the nightly status…

Self-host: verified 28 Sep 2026 · fresh clone of decosa-api on our server, api image built from docker/api/Dockerfile, compose api with a named data volume on the host network, direct route to the running Qwen3.8-27B, page parser and ASR; local signing; torn down after

Measured cost to run: about $0.79 per 100 notes (hosted, 30 Sep 2026). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.

Rehearsal bundle 12/12 four times in a row (8.5-9.6 s); smoke ok in 6.2 s with 17 signed receipts; the photo sample read 9 lines (8 agreed by both readers); 4 synthetic dictations (computer voice, 17-32 s) transcribed and drafted, a 190 s recording refused; /jottings/app drafted a note at 1280 and 390 px with no horizontal scroll; a cold-user test ran a 20-session catch-up against it. Model-server startup was not re-run.

Known limits (6)
  • Hosted verification on production (decosa.ai, gateway route), 28 Sep 2026: the smoke (cbt-panic sample) run 5 times one at a time.
  • Measured on synthetic sessions written by agents and on rendered handwriting; no real jottings, real handwriting or licensed therapist's review yet.
  • Not zero: on 24 new sessions 5 of 326 draft units were still not supported by the jottings (mostly shorthand read the wrong way, such as 'sat' for Saturday).
  • The same jottings can give different drafts on two runs.
  • Dictation was tested with a computer voice only; spoken dates and times are converted to digits, other numbers stay as words.
  • Needs a 32-96 GB GPU in the practice for real notes until a confidential hosted tier with a BAA exists.

Eval results, nightly checks and cost per runVerify a run

How it's builtThe steps, the models and what each one checks
Self-host · your GPUs · recommended

Run it yourself, on request

  • The same open models and app, on 1x RTX PRO 6000 (96 GB) or 1x RTX 5090 (32 GB, estimate) for Qwen3.8-27B; the photo's second reader needs about 3 GB more; typed jottings need nothing else.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Hosted · by Decosa

Get an API key

  • Call the notes from your own jottings API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Synthetic, public or test data only: real patient data belongs on your own hardware.
The open stack

Your own session jottings in, a draft progress note out, every sentence linked to the jotting it came from. No recording, ever.

For therapists in private practice with a notes backlog. After the session, give it what you already jot: typed bullets, a photo of your handwritten notes, or a dictation of up to a minute. Qwen3.8-27B drafts a DAP, SOAP or BIRP progress note tied to your treatment-plan goals, with each sentence citing the jottings it came from. Before a sentence stays, code checks that every number, family member, diagnosis, mood, risk, progress or advice word in it is in those jottings and that jokes, question marks and 'no intent' survive, and the grounding judge must call it supported; a sentence that fails gets one rewrite, then it is removed and listed. Times, dates and the risk line come from code and your own words. Every risk or safety statement you wrote (suicidal or homicidal thoughts, including "denies SI", self-harm, safety plans, access to means, possible reports, substance-use risk) is quoted word for word, any drafted sentence about one uses your own words for it (a paraphrase is replaced by your jotting, quoted), and a draft that would lose one is blocked with the reason. A completeness check lists what payers audit (date, start and stop times, modality, interventions, goals, response, plan, your credential, and risk when you mentioned it) and marks what is missing as [not in your notes]. Catch-up mode drafts up to 30 sessions at once. It never records a session, never diagnoses, assesses risk or recommends treatment, and the signature line stays blank: a draft you review, edit and sign.

Deployment
Self-host first
Regulatory
Checked 28 Sep 2026. Illinois HB 1806, the Wellness and Oversight for Psychological Resources Act (225 ILCS 155; signed 4 Aug 2025, effective on becoming law), allows a licensed professional to use AI for administrative or supplementary support, which includes 'preparing and maintaining client records, including therapy notes' (Sec. 10); requires written informed consent when the client's session is recorded or transcribed (Sec. 15(b)), which this tool never does; and forbids letting AI make independent therapeutic decisions, communicate therapeutically with clients, generate recommendations or treatment plans without review, or detect emotions or mental states (Sec. 20(b)); penalty up to $10,000 per violation. Whether a post-session dictation by the clinician counts as transcribing the session is our reading, not settled: ask counsel. The element list follows CGS's LCD L34353 fact sheet (revised 26 May 2021) and First Coast's LCD L33252 and article A57520 (rev. 1 Jan 2025). Therapy notes are PHI under HIPAA; a cloud service that holds encrypted PHI is still a business associate, so real notes run self-hosted until Decosa has BAAs, and the hosted demo takes made-up sessions only. Not legal advice.
Architecture
Text description

Jottings (typed, a photo, or a dictation of up to 60 seconds made after the session) are read: Qwen3.8-27B reads a photo's page and PaddleOCR-VL-1.6 reads each line again; Qwen3-ASR-1.7B transcribes a dictation. A completeness check (dates, times and risk in code, other elements tagged by the model and gated by code) and the draft (Qwen3.8-27B, each sentence citing jottings) run side by side. Every drafted sentence passes code checks and the grounding judge, or gets one rewrite and is otherwise removed. Out come the draft note with cites, the missing list, the removed list, text to paste into an EHR and a signed record with no note text. Self-hosted, everything stays on the practice's machine; hosted, every model call gets a gateway-signed receipt.

Architecture

At a glance

Data retention
Nothing stored: the jottings, the photo or the dictation live in memory for the request. The signed record holds hashes of each jotting and sentence and receipt ids, never text; logs carry counts only.
What leaves the box
Hosted (made-up sessions only): every model call goes through our gateway to the GPU serving Qwen3.8-27B. Self-hosted on the direct route: nothing leaves the machine.
What it will not do
Record or take a session recording; drop, change or paraphrase a risk or safety statement you wrote (a sentence about one quotes your jotting; a draft that would lose one is blocked); add a diagnosis, a mood or mental state, a risk level or a recommendation; fill in a missing time, date or plan; sign or send anything. Your own ? and rule-out lines are held back for you.
Input formats
Typed jottings (one per line, up to 6,000 characters), a JPEG, PNG or WebP photo of handwritten jottings (up to 8 MB), or a dictation of up to 60 seconds; DAP, SOAP or BIRP; up to 10 treatment-plan goals; catch-up up to 30 typed sessions.
Typical run
One note: a few dozen model calls (each kept sentence gets a grounding check and a meaning check), a fraction of a cent at the gateway list price; under a minute on a busy shared gateway, seconds self-hosted on the direct route. A catch-up: a few minutes.
Pricing basis
A flat self-host licence per clinician (price not set). A cold-user therapist said about $25 a month.
Quality tiers

Pick the tier for the quality you need

Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.

  • In the hosted demo

    Standard

    one GPU for the model (hosted demo)

    Qwen3.8-27B drafts and checks; code checks times, dates, risk, numbers and clinical words. Photos are read by Qwen3.8 alone. This is what the hosted demo runs, on made-up sessions only.

    Models
    • decosa-api jottings (decosa_api/verticals/jottings), importing the grounding judge (vertical 17)
    • Qwen3.8-27B (NVFP4)
    Hardware
    1x RTX PRO 6000 96 GB (measured) or 1x RTX 5090 32 GB (estimate)
    Quality evidence
    • Draft sentences and header lines a blind reviewer found not supported by the jottings (48 held-out synthetic sessions)25 / 701docs/evals/jottings-note.md, test split, 2026-09-28 (blind Claude Code Opus review)
    • The same blind check on 24 new sessions after the fixes the test found5 / 326docs/evals/jottings-note.md, fresh split, 2026-09-28
    • Same reviewer, the same sessions drafted by one plain prompt to the same model (no checks)751 / 1574docs/evals/jottings-note.md, test split
    • Audited elements marked present or missing correctly379 / 384docs/evals/jottings-note.md, test split
    • Elements missing from the jottings that it marked missing56 / 57docs/evals/jottings-note.md, test split
    Latency
    measured on the shared gateway: the first sample in seconds one at a time; under a minute per note in the eval with notes in parallel; seconds on the direct route (self-host style).
    Verification
    Proof: strong
  • In the hosted demo

    Best

    adds the second reader and dictation

    Everything in Standard, plus PaddleOCR-VL-1.6 reading each handwritten line again (lines the readers disagree on are flagged) and Qwen3-ASR-1.7B for a dictation of up to a minute.

    Models
    • decosa-api jottings (decosa_api/verticals/jottings), importing the grounding judge (vertical 17)
    • Qwen3.8-27B (NVFP4)
    • PaddleOCR-VL-1.6 (the document reader's page parser)
    • Qwen3-ASR-1.7B (language pack speech service)
    Hardware
    1x RTX PRO 6000 96 GB: Qwen3.8 plus about 3 GB for the reader and the ASR loaded on demand (measured on our server)
    Quality evidence
    • Handwritten lines read right (12 rendered photos, test split)99 / 99docs/evals/jottings-note.md, test split
    • Lines flagged 'check the reading' that were in fact read right16 of 16 flagsdocs/evals/jottings-note.md, test split (the second reader's disagreements are mostly punctuation or letter case)
    • Dictations of up to 60 s transcribed and drafted; a 190 s recording refused4 / 4; refuseddocs/evals/jottings-note.md, self-host check (synthetic voice)
    Latency
    measured: a photo adds the page read and the line reads; about a minute per note in the test eval on the shared gateway, seconds on the direct route in dev.
    Verification
    Proof: partialThe reader and the ASR are not hosted models: their calls carry the instance's model-call receipts; the model's calls carry gateway receipts when hosted.
Components

Every model in the stack

Models in this stack. Each row has a button that shows its licence, engine, verification and evidence.
ModelDetails
Reading typed jottings, the completeness check (dates, times and risk in code), the number, clinical-claim, qualifier and attribution guards, layouts, catch-up and the signed record (no model; CPU)decosa-api jottings (decosa_api/verticals/jottings), importing the grounding judge (vertical 17)
0 GBProof: partial
Drafts the sentences, tags the audited elements, checks every sentence (the grounding judge), rewrites a failed sentence once, and reads a photo's pageQwen3.8-27B (NVFP4)nvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab)
27.8B · 20 GBProof: strongIn the hosted demo
Second reader for handwriting: reads each line of a photo again, so lines the two readers read differently are flagged for the clinicianPaddleOCR-VL-1.6 (the document reader's page parser)PaddlePaddle/PaddleOCR-VL-1.6 on Hugging Face (opens in a new tab)
0.9B · 3 GBProof: partial
Transcribes a dictation of up to 60 seconds made after the session (never a session recording)Qwen3-ASR-1.7B (language pack speech service)Qwen/Qwen3-ASR-1.7B on Hugging Face (opens in a new tab)
1.7BProof: partial

Around the models

Tools, services and hardware

Tools

Services

  • decosa-api:8445
    ${DECOSA_REGISTRY}/decosa-api:<tag>

    GET /jottings/info, /jottings/samples, /jottings/app (the page for the practice's machine); POST /jottings/note (SSE or JSON), /jottings/batch, /jottings/read; POST /record/verify. Keeps no note text.

  • vLLM (model):8114
    vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1

    Qwen3.8-27B NVFP4 behind our gateway (hosted) or called directly (self-host).

  • vLLM (second reader, optional):8498
    vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1

    PaddleOCR-VL-1.6 for handwritten lines.

Hardware

  • 1x RTX PRO 6000 Blackwell 96 GB Fits

    Measured on our server: the hosted runs and the eval went through the shared gateway; the self-host check ran on the direct route on the same card.

  • 1x RTX 5090 32 GB Fits

    Estimate: Qwen3.8-27B NVFP4 needs about 20 GB of weights plus KV cache; not run for this tool.

Latency per lane

  • one note, the cbt-panic sample, production, one at a time25.5 s

    Measuredmeasured on production 2026-09-30: smoke module, 5 runs, 18.7-30.3 s (p95 30.3 s); the same run as verification.hosted

  • one note, the cbt-panic sample, hosted gateway route, one at a time16.4 s

    Measuredmeasured on our server 2026-09-28: 5 runs, 9.5-16.6 s, pre-release server on the gateway

  • one note, eval, hosted gateway route, 3 in parallel42.4 s

    Measuredmeasured on our server 2026-09-28: p50 over 48 held-out sessions (p95 64.0 s); gateway shared with other workloads

  • one note, direct route (self-host style)9.4 s

    Measuredmeasured on our server 2026-09-28: p50 over the 14 dev sessions

  • catch-up, 20 sessions, hosted gateway route222.3 s

    Measuredmeasured on our server 2026-09-28: one recorded run, 2 sessions in parallel

Notes

  • A blind reviewer (Claude Opus, given only the jottings, the goals and the note) found 25 of 701 sentences and header lines not supported by the jottings on 48 held-out synthetic sessions; for the same sessions drafted by one plain prompt to the same model it found 751 of 1,574, including 454 added clinical claims (mood, affect, diagnoses, progress judgements).
  • After fixing what that test found, 24 new sessions gave 5 of 326 not supported; after the fixes from that run and from a cold-user test, the same 24 sessions gave 2 of 332 (no longer held out).
  • Missing elements: 56 of 57 caught on the test split and 26 of 26 on the fresh one; the fresh split also showed times written on their own line being missed (15 of 166 present elements called missing), fixed afterwards.
  • Handwriting: 99 of 99 lines on 12 rendered photos read right; the second reader flagged 16 lines, all of them read right (false alarms). Rendered fonts, not real handwriting.
  • Everything here is synthetic, written by agents; no licensed therapist has read the outputs yet.
Assemble it

Run this exact stack on your machine

Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.

jottings-note/assemble-prompt.md136 lines
# Assemble "Notes from your own jottings" on this machine

You are setting up a documentation aid for a psychotherapy practice. After a session the therapist gives it their own
jottings (typed bullets, a photo of handwritten notes, or a dictation of up to 60 seconds made after the session) and gets
a draft DAP, SOAP or BIRP progress note where every sentence cites the jotting it came from, plus a list of what an
auditor would find missing. It never takes a session recording. Work step by step, show me each command before you run
anything with `sudo`, and stop to ask if a check fails.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/jottings-note.zip (3 KB, 12 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py jottings-note` (the api image carries the same bundle under /app/rehearsal/jottings-note/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py jottings-note --bundle jottings-note.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "every audited element is in the CBT jottings, so nothing is missing", "the risk line carries the jotting as written", "the start and stop times are the ones written, with the minutes between them"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## 0. Ground rules and licences
- Models: Qwen3.8-27B (Apache-2.0) drafts, checks every sentence and reads photos; PaddleOCR-VL-1.6 (Apache-2.0) is the
  second reader for handwriting (optional); Qwen3-ASR-1.7B (Apache-2.0) transcribes dictations (optional). The pipeline,
  the checks and the signed record are decosa-api (AGPL-3.0-or-later).
- Therapy notes are PHI. They stay on this machine. Bind every port to 127.0.0.1. The service keeps no note text: nothing
  is written to disk and logs carry counts only. Do not add request logging.
- It is a draft for the licensed clinician to review, edit and sign. Do not add features that record sessions, send
  anything to clients, sign notes, or suggest diagnoses or treatment. The tests fail if a recording input appears.

## 1. Check the machine
1. `nvidia-smi`: one GPU with at least 32 GB (Qwen3.8-27B NVFP4 needs about 20 GB of weights plus KV cache; we measured
   on an RTX PRO 6000 96 GB; an RTX 5090 32 GB should fit but we have not run this tool on one). Blackwell runs NVFP4;
   on older cards use `Qwen/Qwen3.8-27B-FP8`. Typed jottings need only this model.
2. `docker --version`, `docker compose version`. If Docker or the NVIDIA container toolkit is missing, install them from
   the official Docker and NVIDIA repositories after asking me, then run
   `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
3. Disk: about 40 GB free (the model, plus about 2 GB for the optional handwriting reader).

## 2. Images and weights
- `${DECOSA_REGISTRY}/decosa-api:<tag>` (**publishing soon**). If the pull fails, build from source:
  `git clone <decosa-api source: on request at https://decosa.ai/contact?topic=self-host>` (access required), check out the newest release tag that
  contains `decosa_api/verticals/jottings/` (`main` until one does), and build `docker/api/Dockerfile`.
- `vllm/vllm-openai:v0.29.0`; weights `nvidia/Qwen3.8-27B-NVFP4` (revision `482ca0f3832238542f8f5295dde86b5f22711d80`).
- Optional, for photos: a second vLLM serving `PaddlePaddle/PaddleOCR-VL-1.6` (revision `c5630ab`) with
  `--gpu-memory-utilization 0.04`. Without it, photos are read by Qwen3.8 alone and nothing is flagged as "readers differ".
- Optional, for dictation: the language pack's speech service (Qwen3-ASR-1.7B), see decosa-api `services/lang/`.

## 3. docker-compose.yml
Write this in `~/decosa/jottings/`:

```yaml
services:
  llm:
    image: vllm/vllm-openai:v0.29.0
    command: ["--model", "nvidia/Qwen3.8-27B-NVFP4", "--served-model-name", "qwen3.8-27b", "--max-model-len", "32768",
              "--enable-prefix-caching"]
    ports: ["127.0.0.1:8114:8000"]
    volumes: ["~/.cache/huggingface:/root/.cache/huggingface"]
    deploy: { resources: { reservations: { devices: [{ driver: nvidia, count: 1, capabilities: [gpu] }] } } }
    healthcheck: { test: ["CMD", "curl", "-fs", "http://localhost:8000/v1/models"], interval: 30s, retries: 20 }
  reader:   # optional: the second reader for handwriting photos
    image: vllm/vllm-openai:v0.29.0
    command: ["--model", "PaddlePaddle/PaddleOCR-VL-1.6", "--revision", "c5630ab", "--served-model-name", "PaddlePaddle/PaddleOCR-VL-1.6",
              "--trust-remote-code", "--gpu-memory-utilization", "0.04", "--max-model-len", "8192", "--max-num-seqs", "16"]
    ports: ["127.0.0.1:8498:8000"]
    volumes: ["~/.cache/huggingface:/root/.cache/huggingface"]
    deploy: { resources: { reservations: { devices: [{ driver: nvidia, count: 1, capabilities: [gpu] }] } } }
  api:
    image: ${DECOSA_REGISTRY}/decosa-api:<tag>
    ports: ["127.0.0.1:8445:8445"]
    environment:
      DECOSA_HOST: 0.0.0.0
      DECOSA_PORT: "8445"
      DECOSA_DATA_DIR: /data
      DECOSA_LLM_ROUTE: direct
      DECOSA_LLM_URL: http://llm:8000/v1
      DECOSA_LLM_MODEL: qwen3.8-27b
      DECOSA_JOTTINGS_SYNTHETIC_ONLY: "0"
      DECOSA_DOCREADER_PARSER_URL: http://reader:8000/v1
      DECOSA_BUDGET_LLM_TOKENS: "400000"
    volumes: ["decosa-data:/data"]
    depends_on: { llm: { condition: service_healthy } }
    healthcheck: { test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://localhost:8445/jottings/info', timeout=4)"], interval: 30s, retries: 10 }
volumes:
  decosa-data:
```

`DECOSA_JOTTINGS_SYNTHETIC_ONLY: "0"` lets this box take real notes; the hosted service takes made-up sessions only and
refuses text that looks like real identifiers. Leave out the `reader` service (and `DECOSA_DOCREADER_PARSER_URL`, or set
`DECOSA_JOTTINGS_SECOND_READER: "0"`) if you only type jottings. The api keeps its state in the named volume
`decosa-data` (a host folder created by Docker is owned by root and stops the unprivileged api). Then
`docker compose up -d`.

On first start the api creates this box's Ed25519 key in `decosa-data` (`/data/attest/`, mode 0600). Back it up with
`docker compose cp api:/data/attest ./attest-backup`, keep it private, never print it. Every model call on the direct
route gets a receipt signed with that key (an attestation by the operator, not a proof). Never set
`DECOSA_LLM_ROUTE=gateway` on this box: that sends PHI off the machine.

## 4. Smoke test
1. `curl -s localhost:8445/jottings/info | jq '{synthetic_only, route: .model.route, elements: [.elements[].id]}'` shows
   `synthetic_only: false`, `direct`, and the ten elements (date ... signature, risk).
2. Token: `T=$(curl -s -XPOST localhost:8445/demo/session -H 'content-type: application/json' -d '{"vertical":"jottings-note"}' | jq -r .token)`.
3. `curl -s -XPOST localhost:8445/jottings/note -H "authorization: Bearer $T" -H 'content-type: application/json' -d '{"sample":"cbt-panic"}' > res.json`,
   then `jq -r .note.plain res.json`. Expect nothing in `.report.missing`, `Total time: 53 minutes`, the risk line
   `"denies SI/HI"` as written, and the mother's opinion kept as her opinion (never "diagnosed with bipolar disorder").
4. `jq '{removed: .report.sentences.removed, missing: .report.missing}' <(curl -s -XPOST localhost:8445/jottings/note -H "authorization: Bearer $T" -H 'content-type: application/json' -d '{"sample":"couples-gaps"}')`:
   `missing` includes `start_time`, `stop_time` and `plan`, and the note says `[not in your notes]` for them.
5. `jq .report.record res.json | curl -s -XPOST localhost:8445/record/verify -H 'content-type: application/json' -d @- | jq .ok`
   must print `true`.
6. A request with a `"recording"` field must return 400 ("never takes a session recording").
7. If decosa-api's source is at hand, `python scripts/rehearse.py jottings-note --base-url http://127.0.0.1:8445` runs the
   checks and prints PASS or FAIL per property (11 checks). On our server it took 49 s.
8. Photos (if the reader runs): `{"sample":"dbt-photo"}` reads nine lines; `.ready.jottings[].read` says `agree` for lines
   both readers read the same.

## 5. Use it
- Open `http://127.0.0.1:8445/jottings/app` in a browser on this machine: type or photograph jottings, pick DAP, SOAP or
  BIRP, add goals and your name and credential, Draft, then Copy the note (or one section at a time) into your EHR. The
  "Catch up on a backlog" tab drafts up to 30 sessions at once.
- Or call `POST /jottings/note` and `POST /jottings/batch` from your own scripts; contract: `API_CONTRACT.md`, section
  "Changes (jottings-note: write a note from my jottings, 2026-09-28)".

Off by default. Never on a box that holds PHI. If I ask for it, follow the provider guide at `/provide` on the site, and
do not enable it without my explicit yes.
Technical detailsModels, where it runs, labels

In short

Last reviewed

What it is
Therapy progress notes from what you already jot after a session: typed bullets, a photo of your handwriting or a one-minute dictation. Every sentence links to the jotting it came from, nothing is added, and what an auditor would miss is marked, never filled in. No session recording, ever.
Who it's for
Therapists in private practice (LPC, LCSW, LMFT, psychologists) who write their own progress notes and are behind on them.
Where it runs
Self-host for real client notes; the hosted demo takes made-up sessions only
Key numbers

On 48 held-out synthetic sessions a blind reviewer found 25 of 701 draft units not supported by the jottings, against 751 of 1,574 for a plain prompt to the same model; synthetic sessions only.

  • 113 / 117 Risk and safety statements carried, fresh blind split #3 (verbatim rule) (held out, n = 117)
  • 7 / 805 Draft units not supported, fresh held-out split (29 Sep, blind review) (held out, n = 805)
  • 0.0072 USD Cost per note at list price (mean), held-out split (held out, n = 60)
  • 25.5 s Median end-to-end run, hosted (QA sweep 2026-09-30)
All results, datasets and caveats
Models
Qwen3.8-27B (reads the photo, drafts, and checks every sentence with the grounding judge) · PaddleOCR-VL-1.6 (second reader for handwriting) · Qwen3-ASR-1.7B (a dictation of up to 60 s)
Where
Self-host for real client notes; the hosted demo takes made-up sessions only
Checks
Every sentence cites its jottings and passes the number, clinical-claim and grounding checks; receipt per model call; signed record of hashes, no note text
Industry
Healthcare
Output
Notes, reports and drafts · Signed record or verdict
Data
Patient data (PHI)
Hardware
1× 96 GB GPU
Licence
Permissive (Apache-2.0, MIT)
Runs in
Self-host
Built from
Signed record

Questions people ask

Does it record my therapy sessions?

No. There is no input for a session recording. You give it your own jottings, a photo of them, or a dictation of up to 60 seconds made after the session; a longer audio file is refused.

Will it add things I didn't write in my therapy progress notes?

It is built not to: every sentence cites the jottings it came from; code checks numbers, family members, diagnosis, mood, risk, progress and advice words against those jottings; a grounding check must call each sentence supported, or it is rewritten once and otherwise removed and listed. On 24 new synthetic sessions a blind reviewer still found 5 of 326 draft units not supported, so you review before signing.

What does it do when something is missing?

It marks it "[not in your notes]" and lists it: date, start and stop times, type of therapy, interventions, goals addressed, response, plan, your name and credential, and risk when you mentioned it. It never fills in a time, a date or a plan.

Can I use it with real client notes?

Yes, on your own machine (self-host), where nothing leaves the box. The hosted demo takes made-up sessions only, because therapy notes are protected health information and Decosa does not have BAAs yet.

Is it allowed under Illinois HB 1806?

That law allows AI for supplementary support, including preparing therapy notes, and requires written consent when a session is recorded or transcribed; this tool never records a session. Whether a clinician's own post-session dictation counts as transcribing the session is not settled: ask your counsel. Not legal advice.

Ask a question or leave feedbackWe read every message and publish useful answers
Questions & feedback

Ask about Notes from your own jottings

We read every message. Questions, comments and our answers show here once we have reviewed and approved them.

Loading questions…

This is a

Plain text. Please leave out personal, patient or client data.

Shown with your message if we publish it. Leave blank to post as “A visitor”.

Nothing appears here until we have read and approved it.