Build a medical chronology
A dated chronology where every line cites the file, page and spot on the page, with conflicting dates, gaps in treatment and prior conditions flagged.
Built on: Document reader, Evidence retrieval, Typed judgment, Signed record
Loading the tool…
Use it your way
Use it from your codeThe hosted API with your key, and prompts to paste into a coding agent
Run it yourself, on request
- The same open models and app, on 1x RTX PRO 6000 (96 GB) for Qwen3.8-27B, plus about 6 GB for the document reader and 10 GB for the embedder and reranker; the checks, merging and record run on CPU.
- Data never leaves your machines, and there are no Decosa charges.
- One prompt for Claude Code or Codex assembles the whole stack.
- Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Get an API key
- Call the medical chronology with page cites API from your own code in minutes.
- Every model answer carries a signed receipt.
- Synthetic, public or test data only: real patient data belongs on your own hardware.
Build with it
Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.
- Base URL
- https://api.decosa.ai
- Auth
Authorization: Bearer $DECOSA_API_KEY(or a demo session token)- Tool id
- medical-chronology
Use the hosted API
# Decosa medical chronology with page cites: use the hosted API
You are wiring Decosa's medical chronology into this project (a case-management tool, a claims or underwriting workflow,
or a record-review vendor's pipeline). Given a packet of medical records (PDFs or images: typed notes, scans and faxes,
handwriting, lab tables) it returns a dated chronology of encounters, diagnoses, procedures, imaging, labs, medications
and referrals, where every line cites the file, page and box it came from; copies merged; conflicting dates, gaps in
treatment and conditions before the date of injury flagged; a Markdown chronology and a signed record. Every model call
has a signed receipt. Use only what is listed below; if you need something else, stop and ask me.
- Base URL: `https://api.decosa.ai`
- Health check: `GET https://api.decosa.ai/healthz`.
- **The hosted API takes synthetic records only.** Real medical records are protected health information (HIPAA): the
hosted demo refuses anything without `"synthetic": true` and must never receive a real patient's records. For real
records, self-host (the "Run it yourself" prompt). Build and test the integration against synthetic packets here.
- This is a paralegal and reviewer aid, not medical or legal advice: a person checks each line at its cited box.
## Auth: API key (or a demo session)
1. Preferred: an API key (`dk_…`) from "Get an API key" on the tool page. Keep it in `DECOSA_API_KEY`, never in code.
Send `Authorization: Bearer $DECOSA_API_KEY`.
2. Without a key: `POST https://api.decosa.ai/demo/session` with `{"vertical": "medical-chronology"}` returns
`{"token", "expires_at", "budget"}`. A demo token runs one packet at a time (409 otherwise); over a limit you get 429
with `Retry-After`.
3. A run needs about 350 generated tokens a page, plus 1,200, left in the budget before it starts (402 otherwise).
## Build a chronology
- `POST /chronology/run` (token). Body: `{"files": [{"name", "file_b64", "media_type"?}], "synthetic": true,
"incident_date"?: "YYYY-MM-DD", "claimed"?: ["low back", ...], "gap_days"?: 14-365 (default 60), "title"?, "stream"?}`
or `{"sample": "whitlock"}`. Limits: 12 files, 12 MB each, 40 pages in all (hosted).
- The JSON answer has `entries` (each with `id`, `date`, `type`, `label`, `status`, `date_check`, `flags`, `conflicts`,
and `cites`: `file`, `page`, `bbox` [x, y, w, h] in 144-DPI page pixels, `element`, `quote`, `role` source / copy /
mention, `chunk`, `byte_start`, `byte_end`), `gaps`, `preexisting` (with `related` from a typed judgment), `copies`,
`counts`, `complete` and `incomplete_pages`, `index` (index and layout hashes), `documents` (each file's document-reader
receipt id), `search_receipts`, `judgments`, `markdown`, `record` and `receipts`.
- With `Accept: text/event-stream` (or `"stream": true`) the events are `accepted`, `ready`, a `read` per page, `index`,
`copies`, a `page_events` per extracted page, `gap` and `conflict`, `report`, `budget` and `done`.
- `POST /record/verify` (no token) `{"record": {...}}` re-checks the signed record. `GET /chronology/info`,
`GET /chronology/samples` and `GET /attest/signing-key` need no token.
## Example (Python, `pip install httpx`)
```python
import base64, httpx, os, pathlib
API = "https://api.decosa.ai"
H = {"Authorization": f"Bearer {os.environ['DECOSA_API_KEY']}"}
files = [{"name": p.name, "file_b64": base64.b64encode(p.read_bytes()).decode()} for p in sorted(pathlib.Path("synthetic_packet").glob("*.pdf"))]
r = httpx.post(f"{API}/chronology/run", headers=H, timeout=900,
json={"files": files, "synthetic": True, "incident_date": "2024-03-08", "claimed": ["low back"]})
r.raise_for_status()
js = r.json()
for e in js["entries"]:
print(e["date"], e["type"], e["label"][:50], e["cites"][0]["cite"], ",".join(e["flags"]))
print("gaps:", [(g["from"], g["to"]) for g in js["gaps"]], "complete:", js["complete"])
pathlib.Path("chronology.md").write_text(js["markdown"])
```
Run it yourself (containers)
On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.
# Decosa medical chronology with page cites: run it yourself (containers)
You are setting up Decosa's medical chronology on this machine, so medical records (protected health information) never
leave it. It reads every page of a record packet with the document reader, indexes the pages with their boxes, has an
open model list the events on each page, keeps only events whose words and date are on the page, merges copies, flags
conflicting dates, gaps in treatment and conditions before the date of injury, and writes a chronology where every line
cites file, page and box, plus a signed record. Nothing is sent to Decosa's hosted API. It is a paralegal and reviewer
aid, not medical or legal advice.
Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.
Ask me before any command that needs sudo, and show me the command first.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/medical-chronology.zip (943 KB, 12 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py medical-chronology` (the api image carries the same bundle under /app/rehearsal/medical-chronology/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py medical-chronology --bundle medical-chronology.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "every page was read and gave an answer", "the rotator cuff repair is listed and its conflicting date is flagged", "the 138-day gap in treatment is found"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
instructions for this distribution (docs.docker.com/engine/install). Install the NVIDIA container toolkit and check
`docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`. One GPU with at least 48 GB.
2. Fetch the compose file:
`mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
Read it. Keep the `llm` service (Qwen3.8-27B on vLLM) and the `api` service. For the `api` service set
`DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b`,
`DECOSA_CHRONO_SYNTHETIC_ONLY=0` and `DECOSA_DOCREADER_SYNTHETIC_ONLY=0` (so this box accepts real records),
`DECOSA_CHRONO_MAX_PAGES=600`, and bind every port to 127.0.0.1. Never set the gateway route on this box.
3. Document reader: add the `parser` (PaddleOCR-VL-1.6 on vLLM, about 4 GB) and `docreader` (Docling layout,
`services/docreader` in the decosa-api source) services as in the tool's assemble prompt, and set
`DECOSA_DOCREADER_URL=http://docreader:8497` on the api.
4. Retrieval: add the `retrieval` service (`services/retrieval`, Qwen3-Embedding-0.6B and Qwen3-Reranker-4B, about 10 GB)
and set `DECOSA_RETRIEVAL_URL=http://retrieval:8499`, or set `DECOSA_CHRONO_RETRIEVAL=bm25` to run without it.
5. Pull and start: `docker compose pull && docker compose up -d --build`. Wait for every health check (the first start
downloads about 25 GB of weights).
6. Check: `curl -fsS http://127.0.0.1:<PORT>/chronology/info` shows `synthetic_only: false`, the document reader
reachable and the retrieval mode; `GET /attest/signing-key` shows this box's public key.
7. Smoke test: get a token with `POST /demo/session {"vertical":"medical-chronology"}` and send `{"sample": "whitlock"}`
to `POST /chronology/run` (16 synthetic pages). Expect the rotator cuff repair on 2024-04-28 flagged `conflict` with
the other date 2024-05-05, one gap from 2024-05-12 to 2024-09-27, the 2022-11-15 low back pain flagged before the
injury, F1 p4 and F5 p2 listed as copies, and a cite with a box on every entry. Send the record to
`POST /record/verify`: `ok` must be true.
8. Report back: the public key, the entry counts, the flags, and how long the run took.
Off by default. Joining as a provider serves other people's requests on this GPU; never do it on a box that holds
patient records. If I ask for it later, follow the Provide page instead of improvising.
Run it on your own hardwareWhat it needs, and the prompt that sets it up
Run it on your own GPU
Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.
Hardware check
Check your own hardware- CPU only, 64 GB RAMDoesn't fit
Qwen3.8-27B (NVFP4) needs a GPU.
- GeForce RTX 4090Doesn't fit
Needs about 34.4 GB of GPU memory at the smallest settings; 24 GB available.
- GeForce RTX 5090Doesn't fit
Needs about 42.4 GB of GPU memory at the smallest settings; 32 GB available.
- 2x GeForce RTX 5090standard tierRuns
The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
- L40Sstandard tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.
- H100 80 GB (SXM)standard tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.
- RTX PRO 6000 Blackwell 96 GBstandard tierRuns
The standard tier fits (72 of 96 GB).
- 2x RTX PRO 6000 Blackwell 96 GBstandard tierRuns
The standard tier fits (72 of 192 GB).
- Apple M3 Ultra (Mac Studio), 96 GBCan't tell
Memory not known for Docling 2.130 with the Heron layout model (document reader block) has no mapped Apple Silicon build; PaddleOCR-VL-1.6 (0.9B, document reader block) has no mapped Apple Silicon build; Evidence retrieval block: Qwen3-Embedding-0.6B + Qwen3-Reranker-4B (decosa-retrieval service) has no mapped Apple Silicon build.
- Apple M5 Max, 64 GBCan't tell
Memory not known for Docling 2.130 with the Heron layout model (document reader block) has no mapped Apple Silicon build; PaddleOCR-VL-1.6 (0.9B, document reader block) has no mapped Apple Silicon build; Evidence retrieval block: Qwen3-Embedding-0.6B + Qwen3-Reranker-4B (decosa-retrieval service) has no mapped Apple Silicon build.
Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.
On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.
- 1
Check the GPU, Docker and the NVIDIA Container Toolkit
The driver must see the GPU, and Docker must be able to pass it into a container.
nvidia-smi docker compose version docker run --rm --gpus all ubuntu nvidia-smi
- 2
Fetch the compose file
One file describes the API and the language model as services.
mkdir -p ~/decosa && cd ~/decosa curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml - 3
Pull and start
The first start downloads pinned model weights, tens of gigabytes.
docker compose pull docker compose up -d
- 4
Check health
Wait until the API reports ok with the language model loaded. Then point your app at the local base URL.
curl -fsS http://localhost:<PORT>/healthz # {"ok": true, "llm": true, ...} curl -fsS -X POST http://localhost:<PORT>/demo/session \ -H 'Content-Type: application/json' -d '{"vertical":"medical-chronology"}'
Set up with a coding agent, rehearse on mock data, then go private
- Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
- Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's
expected.json. Every check must print PASS. - Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
docker compose exec api python scripts/rehearse.py medical-chronology
Download the mock-data bundle (943 KB, 12 checks)expected.json
Dana Whitlock is fictional. After a rear-end collision on 8 March 2024, five providers send records: typed PCP notes (born-digital PDF), a scanned ED note and lab table, a handwritten physical-therapy evaluation and flow sheet, two MRI reports, and an orthopedic file with a scanned operative note. The ED note also arrives as a fax inside the PCP's file, and the lumbar MRI report as a fax inside the orthopedic file. Planted: the rotator cuff repair's date is given differently in the post-procedure note, therapy stops for 138 days, and a 2022 low-back visit comes before the injury. The run must list the events with a page and box on every line, flag the conflict, the gap and the pre-existing low-back condition (as related), merge the faxed copies, and sign a record that verifies and fails when changed. It takes a few minutes: every page is read by the document reader and one model call is made per page.
What the rehearsal checks
- every page was read and gave an answer
- the rotator cuff repair is listed and its conflicting date is flagged
- the 138-day gap in treatment is found
- the 2022 low back pain is flagged before the injury and judged related
- the faxed ED note inside the PCP's file is merged as a copy
- at least 25 entries, each located in the retrieval index
- every cite's quote is located in an index chunk
- each file has a signed document-reader receipt
- the chronology says it is not medical or legal advice
- the signed record verifies
- the record fails once its entry count is changed
- every model call has a signed receipt
Licence: Synthetic: the patient, providers, dates and records are invented (decosa_api/verticals/chronology/synth.py; handwriting in the OFL fonts Kalam and Patrick Hand). Part of decosa-api, AGPL-3.0-or-later.
Prompt for your coding agent
# Decosa medical chronology with page cites: run it yourself (containers)
You are setting up Decosa's medical chronology on this machine, so medical records (protected health information) never
leave it. It reads every page of a record packet with the document reader, indexes the pages with their boxes, has an
open model list the events on each page, keeps only events whose words and date are on the page, merges copies, flags
conflicting dates, gaps in treatment and conditions before the date of injury, and writes a chronology where every line
cites file, page and box, plus a signed record. Nothing is sent to Decosa's hosted API. It is a paralegal and reviewer
aid, not medical or legal advice.
Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.
Ask me before any command that needs sudo, and show me the command first.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/medical-chronology.zip (943 KB, 12 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py medical-chronology` (the api image carries the same bundle under /app/rehearsal/medical-chronology/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py medical-chronology --bundle medical-chronology.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "every page was read and gave an answer", "the rotator cuff repair is listed and its conflicting date is flagged", "the 138-day gap in treatment is found"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
instructions for this distribution (docs.docker.com/engine/install). Install the NVIDIA container toolkit and check
`docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`. One GPU with at least 48 GB.
2. Fetch the compose file:
`mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
Read it. Keep the `llm` service (Qwen3.8-27B on vLLM) and the `api` service. For the `api` service set
`DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b`,
`DECOSA_CHRONO_SYNTHETIC_ONLY=0` and `DECOSA_DOCREADER_SYNTHETIC_ONLY=0` (so this box accepts real records),
`DECOSA_CHRONO_MAX_PAGES=600`, and bind every port to 127.0.0.1. Never set the gateway route on this box.
3. Document reader: add the `parser` (PaddleOCR-VL-1.6 on vLLM, about 4 GB) and `docreader` (Docling layout,
`services/docreader` in the decosa-api source) services as in the tool's assemble prompt, and set
`DECOSA_DOCREADER_URL=http://docreader:8497` on the api.
4. Retrieval: add the `retrieval` service (`services/retrieval`, Qwen3-Embedding-0.6B and Qwen3-Reranker-4B, about 10 GB)
and set `DECOSA_RETRIEVAL_URL=http://retrieval:8499`, or set `DECOSA_CHRONO_RETRIEVAL=bm25` to run without it.
5. Pull and start: `docker compose pull && docker compose up -d --build`. Wait for every health check (the first start
downloads about 25 GB of weights).
6. Check: `curl -fsS http://127.0.0.1:<PORT>/chronology/info` shows `synthetic_only: false`, the document reader
reachable and the retrieval mode; `GET /attest/signing-key` shows this box's public key.
7. Smoke test: get a token with `POST /demo/session {"vertical":"medical-chronology"}` and send `{"sample": "whitlock"}`
to `POST /chronology/run` (16 synthetic pages). Expect the rotator cuff repair on 2024-04-28 flagged `conflict` with
the other date 2024-05-05, one gap from 2024-05-12 to 2024-09-27, the 2022-11-15 low back pain flagged before the
injury, F1 p4 and F5 p2 listed as copies, and a cite with a box on every entry. Send the record to
`POST /record/verify`: `ok` must be true.
8. Report back: the public key, the entry counts, the flags, and how long the run took.
Off by default. Joining as a provider serves other people's requests on this GPU; never do it on a box that holds
patient records. If I ask for it later, follow the Provide page instead of improvising.
Help me customise for my hardware
Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.
GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page
Doesn't fitMedical chronology with page cites on GeForce RTX 5090
Needs about 42.4 GB of GPU memory at the smallest settings; 32 GB available.
Lite · document reader and the model, no retrieval service: what changesuses estimates
- Needs about 33.4 GB of GPU memory at the smallest settings; 32 GB available.
Memory per component
- The event schema and page checks, copies, mer...: decosa-api medical chronology (decosa_api/verticals/chronology), importing the document reader, evidence retrieval (scans into retrieval), typed judgment and dates blocks. CPU. Runs on CPU (vram_gb 0 in stack.json).
- One call per page: Qwen3.8-27B (NVFP4). ~57.6 GB (at least ~28 GB), weights 21.4 GB (from stack.json). Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)
- Finds and orders the regions of every page: Docling 2.130 with the Heron layout model (document reader block). ~1 GB (from stack.json). vram_gb 1 in stack.json.
- Reads each region of a scanned page: PaddleOCR-VL-1.6 (0.9B, document reader block). ~4.4 GB (from stack.json). vram_gb 4.4 in stack.json.
Expected speed
Not measured.
Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.
Setup prompt for this hardware
The self-host prompt for Medical chronology with page cites, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.
# Set up Medical chronology with page cites on my hardware Fetch https://decosa.ai/prompts/medical-chronology-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied. ## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=medical-chronology) Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4). Quality tier: Lite · document reader and the model, no retrieval service (lite). Fit check: doesn't fit; some memory numbers are estimates, not measurements. First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything. Use these components (the setup below describes the standard tier; change it to match): - The event schema and page checks, copies, mer...: decosa-api medical chronology (decosa_api/verticals/chronology), importing the document reader, evidence retrieval (scans into retrieval), typed judgment and dates blocks, CPU - One call per page: Qwen3.8-27B (NVFP4) (nvidia/Qwen3.8-27B-NVFP4), 57.6 GB - Finds and orders the regions of every page: Docling 2.130 with the Heron layout model (document reader block) (docling-project/docling-layout-heron), 1 GB - Reads each region of a scanned page: PaddleOCR-VL-1.6 (0.9B, document reader block) (PaddlePaddle/PaddleOCR-VL-1.6), 4.4 GB Warning: the fit check says this tier does not fit: Needs about 33.4 GB of GPU memory at the smallest settings; 32 GB available. Tell me before going further. During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed. The stack's own component list and compose layout: https://decosa.ai/prompts/medical-chronology-assemble.md
The proof
How we tested itEval results and end-to-end checks, hosted and self-hosted, with dates
Verified end to end
Hosted: verified 27 Sep 2026 · measured 27 Sep 2026: · p50 25 s · ~$0.004 per run · 8 receipts
Loading the nightly status…
Self-host: verified 27 Sep 2026 · fresh clone into a clean directory, api image from docker/api/Dockerfile, compose api with a named volume, direct route to the local Qwen3.8-27B and the running document reader and retrieval services, local signing; torn down after
Measured cost to run: about $0.18 per 100 pages (hosted, 27 Sep 2026, partly estimated). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.
The rehearsal bundle passed 12/12 in 68 s (the 16-page packet uploaded as files: conflict, gap, pre-existing, copies, cites in the index, reader receipts, record verified and failed when changed; every receipt attested) and the smoke module passed in 17.9 s. The reader and retrieval services were not built from the compose file here: the running ones were used.
Known limits (5)
- Hosted verification ran on the pre-release server (decosa-api the pre-release branch on our server, gateway route): the smoke test on one 4-page file, and the full 16-page sample from the console. The production API gets this tool when the branch merges.
- Measured on synthetic packets written by the same author as the prompts, one template family; not on real records or against a nurse reviewer's chronology.
- When the reader skips a line on a scan, the event on it is missed: 9 of 198 planted events were missed, mostly this way.
- Hosted runs are capped at 40 pages; a page takes about 5 s on the shared gateway.
- Billing records and charges are not extracted.
How it's builtThe steps, the models and what each one checks
Run it yourself, on request
- The same open models and app, on 1x RTX PRO 6000 (96 GB) for Qwen3.8-27B, plus about 6 GB for the document reader and 10 GB for the embedder and reranker; the checks, merging and record run on CPU.
- Data never leaves your machines, and there are no Decosa charges.
- One prompt for Claude Code or Codex assembles the whole stack.
- Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Get an API key
- Call the medical chronology with page cites API from your own code in minutes.
- Every model answer carries a signed receipt.
- Synthetic, public or test data only: real patient data belongs on your own hardware.
A packet of medical records in, a dated chronology out: every line cites the file, page and box it came from.
For personal-injury and med-mal attorneys, life and disability insurers and APS reviewers. Give it the records as PDFs or images: typed notes, scanned and faxed forms, handwriting, lab tables. The document reader reads every page into elements with their boxes (and signs a receipt per file); the evidence retrieval block indexes each page's elements as chunks that keep page and box. Qwen3.8-27B reads each page, with the page image for scans, and lists encounters, diagnoses, procedures, imaging, labs, medications and referrals in a fixed schema. Code keeps an event only if its words are on the page and its date is written there (the dates block), turns ordered studies into referrals, merges copies of the same page from different providers, and a search per procedure finds other pages that give a different date; typed judgments settle the unclear pairs and say whether a condition documented before the injury is related to it. You get a Markdown chronology with conflicts, gaps in treatment and pre-existing conditions flagged, and a signed record of hashes. A paralegal and reviewer aid, not medical or legal advice.
- Deployment
- Self-host first
- Regulatory
- Written 27 Sep 2026. Medical records held by HIPAA covered entities (providers, health plans) and their business associates are protected health information; 45 CFR 160.103 names legal and consulting services among the services that make a vendor a business associate (https://www.law.cornell.edu/cfr/text/45/160.103, read 27 Sep 2026), and 45 CFR Part 164 Subpart E sets the privacy rules (https://www.hhs.gov/hipaa/for-professionals/privacy/index.html). A plaintiff's firm that gets records with the patient's authorization is usually not a business associate, and life and disability income insurance are excepted benefits outside HIPAA's 'health plan' (PHS Act 2791(c)(1), cited by 160.103; the list itself was not re-read today, unverified here), but state medical-privacy laws, protective orders and professional duties still apply to the same records. That is why real records are self-host only and the hosted demo refuses anything not marked synthetic. The tool lists what the records say with their page and box; it does not give medical opinions (causation, standard of care, prognosis) or legal advice, and a paralegal, nurse reviewer or attorney checks each line. Model licences: Apache-2.0 (Qwen3.8-27B, PaddleOCR-VL-1.6, Qwen3-Embedding, Qwen3-Reranker), MIT and Apache-2.0 (Docling and its layout weights).
Text description
A record packet goes to decosa-api. The document reader reads every page into elements with page and box and signs a receipt per file. The retrieval block indexes each page's elements as chunks that keep page and box, with index and layout hashes, and copied pages are found. Qwen3.8-27B reads each page, with the image for scans, and lists typed events. Code keeps events whose quote is on the page and whose date is written there, turns orders into referrals, merges copies and mentions, and searches for other pages that give another date; typed judgments settle unclear pairs. The output is a dated chronology citing file, page and box, with conflicts, gaps and pre-existing conditions flagged, and a signed record of hashes.
At a glance
- Data retention
- Nothing stored: files and page images live in memory for the request (the sample packet's reader output is cached in memory by file hash). The signed record holds hashes and receipt ids, never record text or the patient's name; logs carry counts only.
- What leaves the box
- Self-hosted on the direct route: nothing. The reader, the retrieval service, the model and the checks run on the same machine. Hosted: page text and page images go through the Decosa API, and only synthetic records are accepted.
- What every line carries
- Date, type, label, the quote, and a cite to file, page and box (plus the chunk and byte span in the index). Click-through to the box on the sample packet.
- What it is not
- Medical or legal advice. It does not judge causation, standard of care or damages; a paralegal, nurse reviewer or attorney checks each line at its box.
Pick the tier for the quality you need
Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.
Lite
document reader and the model, no retrieval service
DECOSA_CHRONO_RETRIEVAL=bm25: the same reading, extraction, checks, merging and record; cites are located in a keyword index and the conflict search uses BM25 without the reranker, so about 10 GB less GPU memory.
- Models
- decosa-api medical chronology (decosa_api/verticals/chronology), importing the document reader, evidence retrieval (scans into retrieval), typed judgment and dates blocks
- Qwen3.8-27B (NVFP4)
- Docling 2.130 with the Heron layout model (document reader block)
- PaddleOCR-VL-1.6 (0.9B, document reader block)
- Hardware
- 1x GPU with about 26 GB free for Qwen3.8-27B and the reader
- Quality evidence
- Planted events found, conflicts, gaps with BM25 onlynot measured yetthe eval runner supports it (scripts/chronology_eval.py --bm25); not run on the test split
- Latency
- not measured separately; the reranked search is about a third of the merge step on the hosted run
- Verification
- Proof: strongSelf-host onlyEvery model call is receipted; the searches are BM25 in code, so there are no embed or rerank attestations.
- In the hosted demo
Standard
reader, retrieval service and the model (hosted demo)
The document reader, scans into retrieval with Qwen3-Embedding and Qwen3-Reranker, and Qwen3.8-27B per page with the page image. This is what the hosted demo runs, on synthetic records only.
- Models
- decosa-api medical chronology (decosa_api/verticals/chronology), importing the document reader, evidence retrieval (scans into retrieval), typed judgment and dates blocks
- Qwen3.8-27B (NVFP4)
- Docling 2.130 with the Heron layout model (document reader block)
- PaddleOCR-VL-1.6 (0.9B, document reader block)
- Evidence retrieval block: Qwen3-Embedding-0.6B + Qwen3-Reranker-4B (decosa-retrieval service)
- Hardware
- 1x RTX PRO 6000 96 GB (measured on shared cards)
- Quality evidence
- Planted events found (held-out, 6 packets, 96 pages)189 / 198 (95.5%)decosa-api docs/evals/medical-chronology.md, test split, 27 Sep 2026
- Date right / cite on the right page and box, of those founddates 189 / 189, cites 189 / 189decosa-api docs/evals/medical-chronology.md, test split
- Lines that are not planted events2 / 224 (0.9%)decosa-api docs/evals/medical-chronology.md, test split
- Planted date conflicts / gaps / pre-existing flagged6 / 6, 6 / 6, 12 / 12; 0 false flags of each kinddecosa-api docs/evals/medical-chronology.md, test split
- Events on faxed copies carrying the copy's cite57 / 66decosa-api docs/evals/medical-chronology.md, test split
- Latency
- measured: several minutes per hundred pages on the shared gateway route (held-out eval); a couple of minutes for the sample.
- Verification
- Proof: strongEvery Qwen3.8 call is a separate gateway call with a gateway-signed receipt; each file has a signed document-reader receipt; each search has a signed search receipt with the index hash; the record lists them all.
Every model in the stack
| Model | Tiers | Params · VRAM | Verification | Details |
|---|---|---|---|---|
The event schema and page checks, copies, merging, conflicts, gaps, pre-existing, cites to chunks and bytes, the Markdown chronology and the signed record (no model; CPU)decosa-api medical chronology (decosa_api/verticals/chronology), importing the document reader, evidence retrieval (scans into retrieval), typed judgment and dates blocks 0 GBProof: partial | LiteStandard | 0 GB | Proof: partial | |
| ||||
One call per page (the page's elements, plus the page image for a scan) listing typed events with quote and date; typed judgments for duplicate and conflict pairs and for pre-existing conditions; re-reads of regions the parser was unsure ofQwen3.8-27B (NVFP4)nvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab) 27.8B · 20 GBProof: strongIn the hosted demo | LiteStandard | 27.8B · 20 GB | Proof: strongIn the hosted demo | |
| ||||
Finds and orders the regions of every page (text, headings, tables, checkboxes) with their boxesDocling 2.130 with the Heron layout model (document reader block)docling-project/docling-layout-heron on Hugging Face (opens in a new tab) 1 GBProof: partialIn the hosted demo | LiteStandard | 1 GB | Proof: partialIn the hosted demo | |
| ||||
Reads each region of a scanned page (text, handwriting, tables as cells)PaddleOCR-VL-1.6 (0.9B, document reader block)PaddlePaddle/PaddleOCR-VL-1.6 on Hugging Face (opens in a new tab) 0.9B · 4.4 GBProof: partialIn the hosted demo | LiteStandard | 0.9B · 4.4 GB | Proof: partialIn the hosted demo | |
| ||||
Scans into retrieval: each page's elements as chunks with page and box, the index and layout hashes, every cite located to a chunk and byte span, and a reranked search per procedure for other pages that give another dateEvidence retrieval block: Qwen3-Embedding-0.6B + Qwen3-Reranker-4B (decosa-retrieval service)Qwen/Qwen3-Reranker-4B on Hugging Face (opens in a new tab) 0.6B + 4B · 9 GBProof: partialIn the hosted demo | Standard | 0.6B + 4B · 9 GB | Proof: partialIn the hosted demo | |
| ||||
Tools, services and hardware
Tools
- 45 CFR 160.103 (HIPAA definitions) (opens in a new tab)US government work
Business associate and health plan definitions, cited in the regulatory note (read 27 Sep 2026).
- Kalam and Patrick Hand fonts (opens in a new tab)SIL Open Font License 1.1
The handwriting on the synthetic PT forms and flow sheets (eval and sample only).
Services
- decosa-api:8445
${DECOSA_REGISTRY}/decosa-api:<tag>GET /chronology/info, /chronology/samples; POST /chronology/run (SSE or JSON); POST /retrieval/index-scans; POST /record/verify. Keeps no record text.
- decosa-docreader (+ PaddleOCR-VL on vLLM):8497
built from services/docreader; parser on vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1Page image in, regions with boxes out; nothing leaves the box.
- decosa-retrieval:8499
built from services/retrieval (no published image yet)Embedder and reranker for the evidence retrieval block, on the same GPU box.
- vLLM (model):8114
vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1Qwen3.8-27B NVFP4 behind our gateway (hosted) or called directly (self-host).
Hardware
- 1x RTX PRO 6000 Blackwell 96 GB Fits
Measured on our server: the model through the shared gateway on one card; the document reader (about 5.7 GB) and the retrieval service (about 10 GB) on another shared card.
- 1x 48 GB card Fits
Estimate: Qwen3.8-27B NVFP4 (about 20 GB with a small KV cache) plus about 16 GB for the reader and retrieval models. Not measured.
- CPU only Does not fit
The model and the page parser need a GPU. The checks, merging and record run on CPU.
Latency per lane
- the 16-page sample packet, hosted gateway route, streamed144.4 s
Measuredmeasured on our server 2026-09-27 (pre-release server): reading 51.6 s, extraction 43.0 s, merging, judgments and the conflict search 49.8 s; gateway shared with other workloads
- per 100 pages, held-out eval (6 packets of 16 pages, one at a time)537.7 s
Measuredmeasured on our server 2026-09-27: 537.7 s per 100 pages end to end (64 to 141 s a packet), shared gateway, 4 page calls at a time
- smoke test: one 4-page file25.3 s
Measuredmeasured on our server 2026-09-27, hosted gateway route
Notes
- On six held-out synthetic packets (96 pages: typed, scanned, faxed and handwritten) it found 189 of 198 planted events; every found event had the right date and a cite on the right page and box.
- It flagged all 6 planted date conflicts, all 6 gaps in treatment and all 12 conditions documented before the injury (related or not said right 12 of 12), with no false conflict, gap or pre-existing flag; 2 of 224 lines were not planted events.
- Misses come mostly from the reader skipping a line on a scan (an ED diagnosis, an X-ray on a fax) and from the model leaving out the visit itself on a page.
- The packets were written by the same author as the prompts, in one template family; real records are messier and have not been measured.
Run this exact stack on your machine
Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.
# Assemble Decosa medical chronology with page cites on this machine
You are setting up a self-hosted medical chronology for a law firm, an insurer's claims or underwriting team, or a
record-review vendor. It reads a packet of medical records (typed PDFs, scans and faxes, handwriting, lab tables) and
returns:
- a dated chronology of encounters, diagnoses, procedures, imaging, labs, medications and referrals, every line citing
the file, page and box it came from (and the chunk and byte span in the retrieval index);
- copies from different providers merged; conflicting dates for the same procedure, gaps in treatment after the date of
injury, and conditions documented before it flagged;
- a signed, hash-chained record of hashes: each file and its document-reader receipt, the index and layout hashes, every
entry's cites and quote hashes, the judgments and every model receipt id. No record text.
Work step by step, show me each command before running anything that needs sudo, and stop if a check fails.
**Before anything else, remind me:**
- Medical records are protected health information (HIPAA). On this box nothing leaves the machine: the document reader,
the retrieval models, the language model and the checks all run here. Never point it at a hosted gateway while it
- This is a paralegal and reviewer aid, not medical or legal advice. A person checks each line at its box; it does not
judge causation, standard of care or damages.
Repeat both points in your final summary.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/medical-chronology.zip (943 KB, 12 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py medical-chronology` (the api image carries the same bundle under /app/rehearsal/medical-chronology/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py medical-chronology --bundle medical-chronology.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "every page was read and gave an answer", "the rotator cuff repair is listed and its conflicting date is flagged", "the 138-day gap in treatment is found"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## What you are building
| service | image | model | port |
|---|---|---|---|
| `llm` | `${DECOSA_REGISTRY}/decosa-llm:0.1.0` (vLLM 0.29.0, `vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1`) | `nvidia/Qwen3.8-27B-NVFP4` @ `482ca0f3832238542f8f5295dde86b5f22711d80`, Apache-2.0 | internal 8000 |
| `parser` | the same vLLM image | `PaddlePaddle/PaddleOCR-VL-1.6` @ `c5630abae1d940eafe0697512a0325494b02ab42`, Apache-2.0 | internal 8498 |
| `docreader` | built from `decosa-api/services/docreader` | Docling layout `docling-project/docling-layout-heron` @ `8f39ad3c0b4c58e9c2d2c84a38465abf757272d8` (MIT + Apache-2.0) | internal 8497 |
| `retrieval` | built from `decosa-api/services/retrieval` | `Qwen/Qwen3-Embedding-0.6B` @ `97b0c614be4d77ee51c0cef4e5f07c00f9eb65b3` + `Qwen/Qwen3-Reranker-4B` @ `22e683669bc0f0bd69640a1354a6d0aebcfeede5`, Apache-2.0 | internal 8499 |
| `api` | `${DECOSA_REGISTRY}/decosa-api:0.1.0` (no GPU) | none | `127.0.0.1:8445` |
## 1. Check the GPU, driver and Docker
1. Run `nvidia-smi`. I need one NVIDIA GPU with at least 48 GB: the model about 20 GB plus KV cache, the page parser
and layout model about 6 GB, the retrieval models about 10 GB. Driver 580 or newer.
- Blackwell (RTX PRO 6000, B200): the defaults below (NVFP4) are the measured setup.
- Hopper (H100/H200) or a 48 GB card: `LLM_MODEL=Qwen/Qwen3.8-27B-FP8`, `LLM_REVISION=main`, `LLM_GPU_UTIL=0.62`;
on 48 GB also `LLM_MAX_LEN=32768` and `PARSER_GPU_UTIL=0.10`. Not measured.
- Under 48 GB: tell me, and use the BM25 option in step 3 (no `retrieval` service).
2. Check `docker --version`, `docker compose version` and `docker run --rm --gpus all ubuntu nvidia-smi`. If Docker or
the NVIDIA Container Toolkit is missing, install them from the official Docker and NVIDIA repositories, run
`sudo nvidia-ctk runtime configure --runtime=docker` and restart Docker.
3. Confirm about 90 GB of free disk.
## 2. Get the images and the source
The images are **on request** while self-host is in early access: ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry (set it as `DECOSA_REGISTRY`), pull access and the compose file. Try `docker pull ${DECOSA_REGISTRY}/decosa-{llm,api}:0.1.0`. If a pull
fails, clone `decosa-api` into `~/decosa-chrono/decosa-api` and build the api with
`docker build -f docker/api/Dockerfile -t ${DECOSA_REGISTRY}/decosa-api:0.1.0 .`. The `docreader` and `retrieval`
services are always built from the clone, so clone it either way. If neither works, stop and tell me.
## 3. Write the compose file
Create `~/decosa-chrono/.env`:
```bash
DECOSA_TAG=0.1.0
DECOSA_GPU=0
LLM_MODEL=nvidia/Qwen3.8-27B-NVFP4
LLM_REVISION=482ca0f3832238542f8f5295dde86b5f22711d80
LLM_MAX_LEN=65536
LLM_GPU_UTIL=0.62
PARSER_GPU_UTIL=0.05
DECOSA_SIGNER_NAME="<who signs these records, e.g. Example LLP records team>"
```
Create `~/decosa-chrono/docker-compose.yml` with exactly these services:
```yaml
name: decosa-chrono
x-gpu: &gpu { resources: { reservations: { devices: [ { driver: nvidia, device_ids: ["${DECOSA_GPU:-0}"], capabilities: [gpu] } ] } } }
x-health: &health { interval: 15s, timeout: 5s, retries: 5 }
services:
llm:
image: ${DECOSA_REGISTRY}/decosa-llm:${DECOSA_TAG}
deploy: *gpu
ipc: host
restart: unless-stopped
volumes: [hf-cache:/root/.cache/huggingface]
command: ["${LLM_MODEL}", "--revision", "${LLM_REVISION}", "--served-model-name", "qwen3.8-27b",
"--max-model-len", "${LLM_MAX_LEN}", "--gpu-memory-utilization", "${LLM_GPU_UTIL}", "--max-num-seqs", "16",
"--kv-cache-dtype", "fp8_e4m3", "--speculative-config", '{"method":"mtp","num_speculative_tokens":3}',
"--seed", "0", "--enable-force-include-usage", "--host", "0.0.0.0", "--port", "8000"]
healthcheck: { <<: *health, test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8000/health', timeout=4)"], start_period: 900s }
parser:
image: vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1
deploy: *gpu
ipc: host
restart: unless-stopped
volumes: [hf-cache:/root/.cache/huggingface]
command: ["PaddlePaddle/PaddleOCR-VL-1.6", "--revision", "c5630abae1d940eafe0697512a0325494b02ab42", "--trust-remote-code",
"--max-model-len", "8192", "--gpu-memory-utilization", "${PARSER_GPU_UTIL}", "--max-num-seqs", "16",
"--no-enable-prefix-caching", "--mm-processor-cache-gb", "0", "--generation-config", "vllm", "--host", "0.0.0.0", "--port", "8498"]
healthcheck: { <<: *health, test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8498/health', timeout=4)"], start_period: 600s }
docreader:
build: { context: ./decosa-api/services/docreader }
deploy: *gpu
restart: unless-stopped
depends_on: { parser: { condition: service_healthy } }
environment:
DOCREADER_PARSER_URL: http://parser:8498/v1
DOCREADER_PARSER_MODEL: PaddlePaddle/PaddleOCR-VL-1.6
DOCLING_DEVICE: cuda
volumes: [hf-cache:/models]
retrieval:
build:
context: ./decosa-api/services/retrieval
dockerfile_inline: |
FROM python:3.12-slim
WORKDIR /app
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt --extra-index-url https://download.pytorch.org/whl/cu130
COPY *.py .
CMD ["python", "server.py"]
deploy: *gpu
restart: unless-stopped
environment: { RETRIEVAL_HOST: 0.0.0.0, RETRIEVAL_PORT: "8499", RETRIEVAL_EMBED: qwen3-emb-0.6b, RETRIEVAL_RERANK: qwen3-rr-4b, HF_HOME: /root/.cache/huggingface }
volumes: [hf-cache:/root/.cache/huggingface]
healthcheck: { <<: *health, test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8499/health', timeout=4)"], start_period: 600s }
api:
image: ${DECOSA_REGISTRY}/decosa-api:${DECOSA_TAG}
restart: unless-stopped
depends_on: { llm: { condition: service_healthy }, docreader: { condition: service_healthy }, retrieval: { condition: service_healthy } }
environment:
DECOSA_LLM_ROUTE: direct # local model only; receipts are signed by this box's key ("attested")
DECOSA_LLM_URL: http://llm:8000/v1
DECOSA_LLM_MODEL: qwen3.8-27b
DECOSA_LOCAL_SIGNING: "on" # Ed25519 key created at /data/attest/ed25519.pem on first start
DECOSA_SIGNER_NAME: ${DECOSA_SIGNER_NAME}
DECOSA_DOCREADER_URL: http://docreader:8497
DECOSA_RETRIEVAL_URL: http://retrieval:8499
DECOSA_CHRONO_SYNTHETIC_ONLY: "0" # this box accepts real records
DECOSA_DOCREADER_SYNTHETIC_ONLY: "0"
DECOSA_CHRONO_MAX_PAGES: "600" # pages per run (hosted demo: 40)
DECOSA_SESSIONS_PER_IP_HOUR: "1000"
DECOSA_BUDGET_LLM_TOKENS: "600000" # per session; about 350 generated tokens a page
DECOSA_SESSION_TTL_S: "28800"
DECOSA_LLM_TIMEOUT_S: "300"
DECOSA_CORS_ORIGIN_REGEX: '^https?://(localhost|127\.0\.0\.1)(:\d+)?$$'
ports: ["127.0.0.1:8445:8445"]
volumes: [decosa-data:/data]
healthcheck: { <<: *health, test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8445/healthz', timeout=4)"], start_period: 20s }
volumes: { hf-cache: {}, decosa-data: {} }
```
**No room for the retrieval models?** Delete `retrieval` and its `depends_on` entry, drop `DECOSA_RETRIEVAL_URL`, and add
`DECOSA_CHRONO_RETRIEVAL: bm25` to the api: cites are still located in the index by keyword chunks; the conflict sweep
uses BM25 instead of the reranker.
Use the named volume `decosa-data` exactly as written: a root-owned host bind mount makes the API fail on
`/data/keys.sqlite`. Run `docker compose up -d --build` and poll `docker compose ps` until every service is healthy (the
LLM takes 5 to 10 minutes the first time). `curl -s localhost:8445/chronology/info | jq '{synthetic_only, blocks}'` should
show `synthetic_only: false`, the document reader reachable and retrieval `hybrid+rerank`.
## 4. Smoke test on the bundled synthetic packet
```bash
API=localhost:8445
TOKEN=$(curl -s $API/demo/session -H 'content-type: application/json' -d '{"vertical":"medical-chronology"}' | jq -r .token)
curl -s $API/chronology/run -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' -d '{"sample":"whitlock"}' > /tmp/chrono.json
jq -r '.entries[] | "\(.date)\t\(.type)\t\(.label[0:50])\t\(.cites[0].cite)\t\(.flags|join(","))"' /tmp/chrono.json
jq '{counts, gaps: [.gaps[] | {from, to, days}], preexisting: [.preexisting[] | {date, label, related}], copies}' /tmp/chrono.json
jq '{record}' /tmp/chrono.json | curl -s $API/record/verify -H 'content-type: application/json' -d @- | jq '{ok, summary}'
```
Pass if: the procedure "right shoulder arthroscopic rotator cuff repair" on 2024-04-28 carries the flag `conflict` with
the other date 2024-05-05; there is one gap from 2024-05-12 to 2024-09-27; the 2022-11-15 low back pain diagnosis is
flagged before the injury and judged related; the pages F1 p4 and F5 p2 are listed as copies; every entry has at least one
cite with a box; and the record verifies. About 20 to 40 entries is normal; the planted answer key has 32 required events.
## 5. Run a real packet
```bash
python3 - <<'PY' > packet.json
import base64, json, pathlib
files = [{"name": p.name, "file_b64": base64.b64encode(p.read_bytes()).decode()} for p in sorted(pathlib.Path("records").iterdir())
if p.suffix.lower() in (".pdf", ".png", ".jpg", ".jpeg", ".tif", ".tiff")]
print(json.dumps({"files": files, "incident_date": "2024-03-08", "claimed": ["low back", "right shoulder"], "gap_days": 60}))
PY
curl -s $API/chronology/run -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' -d @packet.json > run.json
jq -r .markdown run.json > chronology.md
jq .record run.json > chronology-record.json
```
Limits per run: 12 files of 12 MB, and `DECOSA_CHRONO_MAX_PAGES` pages in all. Measured on our server's shared gateway:
about 5 to 7 s a page end to end; a private GPU runs faster. For a very large packet, split it by provider and run
each, or raise the limits. Check `complete` and `incomplete_pages` in every result.
## 6. Point tools at the local API
- Base URL `http://localhost:8445`; for a web app set `NEXT_PUBLIC_DECOSA_API=http://localhost:8445`.
- `POST /chronology/run` returns JSON, or Server-Sent Events with `Accept: text/event-stream`. `POST /docreader/read` and
`POST /retrieval/index-scans` are the two building blocks on their own.
- Keep the chronology and the record with the case file. Anyone can re-check the record with `POST /record/verify`
against the key at `GET /attest/signing-key`.
- Keep the API on 127.0.0.1; for other users, put a TLS reverse proxy with authentication in front, and handle the box
under your HIPAA security program (access control, audit logs, encryption at rest).
## 7. Keep the direct route
`DECOSA_LLM_ROUTE=gateway` would send page text and images to the hosted Decosa
API. Leave it off: this box holds PHI.
Finish with a summary: what is running, the health output, the smoke-test results, and the reminders above.Rules and regulations it checks againstDated, linked to the primary source; not legal advice
Regulation watch
Loading the watch status…
4 laws, rules and guidance pages cited; 3 watched nightly at the primary source. A change marks this page for a human re-check; nothing is edited automatically. What we cite and how it is watched
Technical detailsModels, where it runs, labels
In short
Last reviewed
- What it is
- A medical chronology tool for record reviewers: every line of the chronology cites the file, page and box it came from, and copies, conflicting dates, gaps in treatment and pre-existing conditions are flagged.
- Who it's for
- Personal-injury and med-mal attorneys and paralegals, nurse reviewers, and life and disability insurers reviewing attending physician statements.
- Where it runs
- Self-host for real records (PHI); the hosted demo takes synthetic records only
- Key numbers
On six held-out synthetic packets (96 pages) it found 189 of 198 planted events, each with the right date and a cite on the right page and box, and 2 of 224 lines were not planted events; real records are not measured yet.
- 189 / 198 Planted events found (test split, n = 198)
- 189 / 189 Date right, of events found (test split, n = 189)
- 189 / 189 Cite on the right page and box, of events found (test split, n = 189)
- 25.3 s Median end-to-end run, hosted (QA sweep 2026-09-27)
- Models
- Qwen3.8-27B (events per page with the page image, typed judgments) · document reader (Docling layout + PaddleOCR-VL) · Qwen3-Embedding + Qwen3-Reranker (retrieval over the page boxes)
- Where
- Self-host for real records (PHI); the hosted demo takes synthetic records only
- Checks
- Receipt per model call; a signed document-reader receipt per file; retrieval index and layout hashes with signed search receipts; every quote found on its page; signed hash-chained record
- Industry
- Legal · Healthcare
- Input
- Files and media
- Runs
- Self-host
- Output
- Structured data · Signed record or verdict
- Data
- Patient data (PHI)
- Hardware
- 1× 96 GB GPU
- Licence
- Permissive (Apache-2.0, MIT)
- Runs in
- Self-host
- Built from
- Document reader · Evidence retrieval · Typed judgment · Signed record
Questions people ask
Can it build a medical chronology from scanned and handwritten records?
Yes: every page goes through the document reader, which reads scans, faxes, handwriting and tables and gives each element a box. On synthetic packets with scanned, faxed and handwritten pages it found 189 of 198 planted events. Handwriting is read less reliably than print, so check handwritten dates at the cited box.
Can it make up an event?
An event reaches the chronology only if its quoted words are found on the page it cites and its date is written there; everything else is dropped or flagged. On the held-out packets 2 of 224 lines were not planted events (therapy techniques from a plan listed as procedures).
Do patient records leave our network?
Not when you self-host: the document reader, the retrieval service and the model all run on your own GPU box. The hosted demo accepts synthetic records only and keeps nothing.
Does it find pre-existing conditions and gaps in treatment?
Give it the date of injury and the claimed injuries. It flags diagnoses documented before that date, with a typed judgment on whether each is related, and any gap between visits longer than the gap you set (60 days by default). On the held-out packets it found 12 of 12 pre-existing conditions and 6 of 6 gaps with no false flags.
Is this medical or legal advice?
No. It is a paralegal and reviewer aid: it lists what the records say with the page and box, and a person checks each line. It does not judge causation, standard of care or damages.
Ask a question or leave feedbackWe read every message and publish useful answers
Ask about Medical chronology with page cites
We read every message. Questions, comments and our answers show here once we have reviewed and approved them.
Loading questions…