Skip to content
decosa
LiveHostedSelf-hostSelf-host first for real data

Build a medical chronology

A dated chronology where every line cites the file, page and spot on the page, with conflicting dates, gaps in treatment and prior conditions flagged.

Held-out test189 / 198Planted medical events found (held-out test)
On production25 smedian on production (2026-09-27); slower when the service is busy
List price~$0.18 per 100 pagesmeasured, at list price

Built on: Document reader, Evidence retrieval, Typed judgment, Signed record

Loading the tool…

Use it your way

Use it from your codeThe hosted API with your key, and prompts to paste into a coding agent
Self-host · your GPUs · recommended

Run it yourself, on request

  • The same open models and app, on 1x RTX PRO 6000 (96 GB) for Qwen3.8-27B, plus about 6 GB for the document reader and 10 GB for the embedder and reranker; the checks, merging and record run on CPU.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Hosted · by Decosa

Get an API key

  • Call the medical chronology with page cites API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Synthetic, public or test data only: real patient data belongs on your own hardware.

Build with it

Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.

Base URL
https://api.decosa.ai
Auth
Authorization: Bearer $DECOSA_API_KEY (or a demo session token)
Tool id
medical-chronology

Use the hosted API

# Decosa medical chronology with page cites: use the hosted API

You are wiring Decosa's medical chronology into this project (a case-management tool, a claims or underwriting workflow,
or a record-review vendor's pipeline). Given a packet of medical records (PDFs or images: typed notes, scans and faxes,
handwriting, lab tables) it returns a dated chronology of encounters, diagnoses, procedures, imaging, labs, medications
and referrals, where every line cites the file, page and box it came from; copies merged; conflicting dates, gaps in
treatment and conditions before the date of injury flagged; a Markdown chronology and a signed record. Every model call
has a signed receipt. Use only what is listed below; if you need something else, stop and ask me.

- Base URL: `https://api.decosa.ai`
- Health check: `GET https://api.decosa.ai/healthz`.
- **The hosted API takes synthetic records only.** Real medical records are protected health information (HIPAA): the
  hosted demo refuses anything without `"synthetic": true` and must never receive a real patient's records. For real
  records, self-host (the "Run it yourself" prompt). Build and test the integration against synthetic packets here.
- This is a paralegal and reviewer aid, not medical or legal advice: a person checks each line at its cited box.

## Auth: API key (or a demo session)
1. Preferred: an API key (`dk_…`) from "Get an API key" on the tool page. Keep it in `DECOSA_API_KEY`, never in code.
   Send `Authorization: Bearer $DECOSA_API_KEY`.
2. Without a key: `POST https://api.decosa.ai/demo/session` with `{"vertical": "medical-chronology"}` returns
   `{"token", "expires_at", "budget"}`. A demo token runs one packet at a time (409 otherwise); over a limit you get 429
   with `Retry-After`.
3. A run needs about 350 generated tokens a page, plus 1,200, left in the budget before it starts (402 otherwise).

## Build a chronology
- `POST /chronology/run` (token). Body: `{"files": [{"name", "file_b64", "media_type"?}], "synthetic": true,
  "incident_date"?: "YYYY-MM-DD", "claimed"?: ["low back", ...], "gap_days"?: 14-365 (default 60), "title"?, "stream"?}`
  or `{"sample": "whitlock"}`. Limits: 12 files, 12 MB each, 40 pages in all (hosted).
- The JSON answer has `entries` (each with `id`, `date`, `type`, `label`, `status`, `date_check`, `flags`, `conflicts`,
  and `cites`: `file`, `page`, `bbox` [x, y, w, h] in 144-DPI page pixels, `element`, `quote`, `role` source / copy /
  mention, `chunk`, `byte_start`, `byte_end`), `gaps`, `preexisting` (with `related` from a typed judgment), `copies`,
  `counts`, `complete` and `incomplete_pages`, `index` (index and layout hashes), `documents` (each file's document-reader
  receipt id), `search_receipts`, `judgments`, `markdown`, `record` and `receipts`.
- With `Accept: text/event-stream` (or `"stream": true`) the events are `accepted`, `ready`, a `read` per page, `index`,
  `copies`, a `page_events` per extracted page, `gap` and `conflict`, `report`, `budget` and `done`.
- `POST /record/verify` (no token) `{"record": {...}}` re-checks the signed record. `GET /chronology/info`,
  `GET /chronology/samples` and `GET /attest/signing-key` need no token.

## Example (Python, `pip install httpx`)
```python
import base64, httpx, os, pathlib
API = "https://api.decosa.ai"
H = {"Authorization": f"Bearer {os.environ['DECOSA_API_KEY']}"}
files = [{"name": p.name, "file_b64": base64.b64encode(p.read_bytes()).decode()} for p in sorted(pathlib.Path("synthetic_packet").glob("*.pdf"))]
r = httpx.post(f"{API}/chronology/run", headers=H, timeout=900,
               json={"files": files, "synthetic": True, "incident_date": "2024-03-08", "claimed": ["low back"]})
r.raise_for_status()
js = r.json()
for e in js["entries"]:
    print(e["date"], e["type"], e["label"][:50], e["cites"][0]["cite"], ",".join(e["flags"]))
print("gaps:", [(g["from"], g["to"]) for g in js["gaps"]], "complete:", js["complete"])
pathlib.Path("chronology.md").write_text(js["markdown"])
```

Run it yourself (containers)

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

# Decosa medical chronology with page cites: run it yourself (containers)

You are setting up Decosa's medical chronology on this machine, so medical records (protected health information) never
leave it. It reads every page of a record packet with the document reader, indexes the pages with their boxes, has an
open model list the events on each page, keeps only events whose words and date are on the page, merges copies, flags
conflicting dates, gaps in treatment and conditions before the date of injury, and writes a chronology where every line
cites file, page and box, plus a signed record. Nothing is sent to Decosa's hosted API. It is a paralegal and reviewer
aid, not medical or legal advice.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/medical-chronology.zip (943 KB, 12 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py medical-chronology` (the api image carries the same bundle under /app/rehearsal/medical-chronology/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py medical-chronology --bundle medical-chronology.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "every page was read and gave an answer", "the rotator cuff repair is listed and its conflicting date is flagged", "the 138-day gap in treatment is found"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
   instructions for this distribution (docs.docker.com/engine/install). Install the NVIDIA container toolkit and check
   `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`. One GPU with at least 48 GB.
2. Fetch the compose file:
   `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
   Read it. Keep the `llm` service (Qwen3.8-27B on vLLM) and the `api` service. For the `api` service set
   `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b`,
   `DECOSA_CHRONO_SYNTHETIC_ONLY=0` and `DECOSA_DOCREADER_SYNTHETIC_ONLY=0` (so this box accepts real records),
   `DECOSA_CHRONO_MAX_PAGES=600`, and bind every port to 127.0.0.1. Never set the gateway route on this box.
3. Document reader: add the `parser` (PaddleOCR-VL-1.6 on vLLM, about 4 GB) and `docreader` (Docling layout,
   `services/docreader` in the decosa-api source) services as in the tool's assemble prompt, and set
   `DECOSA_DOCREADER_URL=http://docreader:8497` on the api.
4. Retrieval: add the `retrieval` service (`services/retrieval`, Qwen3-Embedding-0.6B and Qwen3-Reranker-4B, about 10 GB)
   and set `DECOSA_RETRIEVAL_URL=http://retrieval:8499`, or set `DECOSA_CHRONO_RETRIEVAL=bm25` to run without it.
5. Pull and start: `docker compose pull && docker compose up -d --build`. Wait for every health check (the first start
   downloads about 25 GB of weights).
6. Check: `curl -fsS http://127.0.0.1:<PORT>/chronology/info` shows `synthetic_only: false`, the document reader
   reachable and the retrieval mode; `GET /attest/signing-key` shows this box's public key.
7. Smoke test: get a token with `POST /demo/session {"vertical":"medical-chronology"}` and send `{"sample": "whitlock"}`
   to `POST /chronology/run` (16 synthetic pages). Expect the rotator cuff repair on 2024-04-28 flagged `conflict` with
   the other date 2024-05-05, one gap from 2024-05-12 to 2024-09-27, the 2022-11-15 low back pain flagged before the
   injury, F1 p4 and F5 p2 listed as copies, and a cite with a box on every entry. Send the record to
   `POST /record/verify`: `ok` must be true.
8. Report back: the public key, the entry counts, the flags, and how long the run took.

Off by default. Joining as a provider serves other people's requests on this GPU; never do it on a box that holds
patient records. If I ask for it later, follow the Provide page instead of improvising.
Run it on your own hardwareWhat it needs, and the prompt that sets it up

Run it on your own GPU

Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.

  • CPU only, 64 GB RAMDoesn't fit

    Qwen3.8-27B (NVFP4) needs a GPU.

  • GeForce RTX 4090Doesn't fit

    Needs about 34.4 GB of GPU memory at the smallest settings; 24 GB available.

  • GeForce RTX 5090Doesn't fit

    Needs about 42.4 GB of GPU memory at the smallest settings; 32 GB available.

  • 2x GeForce RTX 5090standard tierRuns

    The standard tier fits with changes: Qwen3.8-27B (NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.

  • L40Sstandard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.

  • H100 80 GB (SXM)standard tierRuns

    The standard tier fits with changes: Replace Qwen3.8-27B (NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.

  • RTX PRO 6000 Blackwell 96 GBstandard tierRuns

    The standard tier fits (72 of 96 GB).

  • 2x RTX PRO 6000 Blackwell 96 GBstandard tierRuns

    The standard tier fits (72 of 192 GB).

  • Apple M3 Ultra (Mac Studio), 96 GBCan't tell

    Memory not known for Docling 2.130 with the Heron layout model (document reader block) has no mapped Apple Silicon build; PaddleOCR-VL-1.6 (0.9B, document reader block) has no mapped Apple Silicon build; Evidence retrieval block: Qwen3-Embedding-0.6B + Qwen3-Reranker-4B (decosa-retrieval service) has no mapped Apple Silicon build.

  • Apple M5 Max, 64 GBCan't tell

    Memory not known for Docling 2.130 with the Heron layout model (document reader block) has no mapped Apple Silicon build; PaddleOCR-VL-1.6 (0.9B, document reader block) has no mapped Apple Silicon build; Evidence retrieval block: Qwen3-Embedding-0.6B + Qwen3-Reranker-4B (decosa-retrieval service) has no mapped Apple Silicon build.

Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

  1. 1

    Check the GPU, Docker and the NVIDIA Container Toolkit

    The driver must see the GPU, and Docker must be able to pass it into a container.

    nvidia-smi
    docker compose version
    docker run --rm --gpus all ubuntu nvidia-smi
  2. 2

    Fetch the compose file

    One file describes the API and the language model as services.

    mkdir -p ~/decosa && cd ~/decosa
    curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml
  3. 3

    Pull and start

    The first start downloads pinned model weights, tens of gigabytes.

    docker compose pull
    docker compose up -d
  4. 4

    Check health

    Wait until the API reports ok with the language model loaded. Then point your app at the local base URL.

    curl -fsS http://localhost:<PORT>/healthz
    # {"ok": true, "llm": true, ...}
    curl -fsS -X POST http://localhost:<PORT>/demo/session \
      -H 'Content-Type: application/json' -d '{"vertical":"medical-chronology"}'

Set up with a coding agent, rehearse on mock data, then go private

  1. Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
  2. Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's expected.json. Every check must print PASS.
  3. Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
Rehearsal command
docker compose exec api python scripts/rehearse.py medical-chronology

Download the mock-data bundle (943 KB, 12 checks)expected.json

Dana Whitlock is fictional. After a rear-end collision on 8 March 2024, five providers send records: typed PCP notes (born-digital PDF), a scanned ED note and lab table, a handwritten physical-therapy evaluation and flow sheet, two MRI reports, and an orthopedic file with a scanned operative note. The ED note also arrives as a fax inside the PCP's file, and the lumbar MRI report as a fax inside the orthopedic file. Planted: the rotator cuff repair's date is given differently in the post-procedure note, therapy stops for 138 days, and a 2022 low-back visit comes before the injury. The run must list the events with a page and box on every line, flag the conflict, the gap and the pre-existing low-back condition (as related), merge the faxed copies, and sign a record that verifies and fails when changed. It takes a few minutes: every page is read by the document reader and one model call is made per page.

What the rehearsal checks
  • every page was read and gave an answer
  • the rotator cuff repair is listed and its conflicting date is flagged
  • the 138-day gap in treatment is found
  • the 2022 low back pain is flagged before the injury and judged related
  • the faxed ED note inside the PCP's file is merged as a copy
  • at least 25 entries, each located in the retrieval index
  • every cite's quote is located in an index chunk
  • each file has a signed document-reader receipt
  • the chronology says it is not medical or legal advice
  • the signed record verifies
  • the record fails once its entry count is changed
  • every model call has a signed receipt

Licence: Synthetic: the patient, providers, dates and records are invented (decosa_api/verticals/chronology/synth.py; handwriting in the OFL fonts Kalam and Patrick Hand). Part of decosa-api, AGPL-3.0-or-later.

Prompt for your coding agent

# Decosa medical chronology with page cites: run it yourself (containers)

You are setting up Decosa's medical chronology on this machine, so medical records (protected health information) never
leave it. It reads every page of a record packet with the document reader, indexes the pages with their boxes, has an
open model list the events on each page, keeps only events whose words and date are on the page, merges copies, flags
conflicting dates, gaps in treatment and conditions before the date of injury, and writes a chronology where every line
cites file, page and box, plus a signed record. Nothing is sent to Decosa's hosted API. It is a paralegal and reviewer
aid, not medical or legal advice.

Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with
"not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.

Ask me before any command that needs sudo, and show me the command first.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/medical-chronology.zip (943 KB, 12 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py medical-chronology` (the api image carries the same bundle under /app/rehearsal/medical-chronology/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py medical-chronology --bundle medical-chronology.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "every page was read and gave an answer", "the rotator cuff repair is listed and its conflicting date is flagged", "the 138-day gap in treatment is found"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin using Docker's official
   instructions for this distribution (docs.docker.com/engine/install). Install the NVIDIA container toolkit and check
   `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`. One GPU with at least 48 GB.
2. Fetch the compose file:
   `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`
   Read it. Keep the `llm` service (Qwen3.8-27B on vLLM) and the `api` service. For the `api` service set
   `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b`,
   `DECOSA_CHRONO_SYNTHETIC_ONLY=0` and `DECOSA_DOCREADER_SYNTHETIC_ONLY=0` (so this box accepts real records),
   `DECOSA_CHRONO_MAX_PAGES=600`, and bind every port to 127.0.0.1. Never set the gateway route on this box.
3. Document reader: add the `parser` (PaddleOCR-VL-1.6 on vLLM, about 4 GB) and `docreader` (Docling layout,
   `services/docreader` in the decosa-api source) services as in the tool's assemble prompt, and set
   `DECOSA_DOCREADER_URL=http://docreader:8497` on the api.
4. Retrieval: add the `retrieval` service (`services/retrieval`, Qwen3-Embedding-0.6B and Qwen3-Reranker-4B, about 10 GB)
   and set `DECOSA_RETRIEVAL_URL=http://retrieval:8499`, or set `DECOSA_CHRONO_RETRIEVAL=bm25` to run without it.
5. Pull and start: `docker compose pull && docker compose up -d --build`. Wait for every health check (the first start
   downloads about 25 GB of weights).
6. Check: `curl -fsS http://127.0.0.1:<PORT>/chronology/info` shows `synthetic_only: false`, the document reader
   reachable and the retrieval mode; `GET /attest/signing-key` shows this box's public key.
7. Smoke test: get a token with `POST /demo/session {"vertical":"medical-chronology"}` and send `{"sample": "whitlock"}`
   to `POST /chronology/run` (16 synthetic pages). Expect the rotator cuff repair on 2024-04-28 flagged `conflict` with
   the other date 2024-05-05, one gap from 2024-05-12 to 2024-09-27, the 2022-11-15 low back pain flagged before the
   injury, F1 p4 and F5 p2 listed as copies, and a cite with a box on every entry. Send the record to
   `POST /record/verify`: `ok` must be true.
8. Report back: the public key, the entry counts, the flags, and how long the run took.

Off by default. Joining as a provider serves other people's requests on this GPU; never do it on a box that holds
patient records. If I ask for it later, follow the Provide page instead of improvising.

Help me customise for my hardware

Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.

Hardware

GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page

Doesn't fitMedical chronology with page cites on GeForce RTX 5090

Needs about 42.4 GB of GPU memory at the smallest settings; 32 GB available.

Lite · document reader and the model, no retrieval service: what changesuses estimates

  • Needs about 33.4 GB of GPU memory at the smallest settings; 32 GB available.
Memory per component
  • The event schema and page checks, copies, mer...: decosa-api medical chronology (decosa_api/verticals/chronology), importing the document reader, evidence retrieval (scans into retrieval), typed judgment and dates blocks. CPU. Runs on CPU (vram_gb 0 in stack.json).
  • One call per page: Qwen3.8-27B (NVFP4). ~57.6 GB (at least ~28 GB), weights 21.4 GB (from stack.json). Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 20 GB for this component.)
  • Finds and orders the regions of every page: Docling 2.130 with the Heron layout model (document reader block). ~1 GB (from stack.json). vram_gb 1 in stack.json.
  • Reads each region of a scanned page: PaddleOCR-VL-1.6 (0.9B, document reader block). ~4.4 GB (from stack.json). vram_gb 4.4 in stack.json.

Expected speed

Not measured.

Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.

Setup prompt for this hardware

The self-host prompt for Medical chronology with page cites, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.

# Set up Medical chronology with page cites on my hardware

Fetch https://decosa.ai/prompts/medical-chronology-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied.

## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=medical-chronology)

Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4).
Quality tier: Lite · document reader and the model, no retrieval service (lite). Fit check: doesn't fit; some memory numbers are estimates, not measurements.

First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything.

Use these components (the setup below describes the standard tier; change it to match):
- The event schema and page checks, copies, mer...: decosa-api medical chronology (decosa_api/verticals/chronology), importing the document reader, evidence retrieval (scans into retrieval), typed judgment and dates blocks, CPU
- One call per page: Qwen3.8-27B (NVFP4) (nvidia/Qwen3.8-27B-NVFP4), 57.6 GB
- Finds and orders the regions of every page: Docling 2.130 with the Heron layout model (document reader block) (docling-project/docling-layout-heron), 1 GB
- Reads each region of a scanned page: PaddleOCR-VL-1.6 (0.9B, document reader block) (PaddlePaddle/PaddleOCR-VL-1.6), 4.4 GB

Warning: the fit check says this tier does not fit: Needs about 33.4 GB of GPU memory at the smallest settings; 32 GB available. Tell me before going further.

During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed.

The stack's own component list and compose layout: https://decosa.ai/prompts/medical-chronology-assemble.md

The proof

How we tested itEval results and end-to-end checks, hosted and self-hosted, with dates

Verified end to end

Hosted: verified 27 Sep 2026 · measured 27 Sep 2026: · p50 25 s · ~$0.004 per run · 8 receipts

Loading the nightly status…

Self-host: verified 27 Sep 2026 · fresh clone into a clean directory, api image from docker/api/Dockerfile, compose api with a named volume, direct route to the local Qwen3.8-27B and the running document reader and retrieval services, local signing; torn down after

Measured cost to run: about $0.18 per 100 pages (hosted, 27 Sep 2026, partly estimated). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.

The rehearsal bundle passed 12/12 in 68 s (the 16-page packet uploaded as files: conflict, gap, pre-existing, copies, cites in the index, reader receipts, record verified and failed when changed; every receipt attested) and the smoke module passed in 17.9 s. The reader and retrieval services were not built from the compose file here: the running ones were used.

Known limits (5)
  • Hosted verification ran on the pre-release server (decosa-api the pre-release branch on our server, gateway route): the smoke test on one 4-page file, and the full 16-page sample from the console. The production API gets this tool when the branch merges.
  • Measured on synthetic packets written by the same author as the prompts, one template family; not on real records or against a nurse reviewer's chronology.
  • When the reader skips a line on a scan, the event on it is missed: 9 of 198 planted events were missed, mostly this way.
  • Hosted runs are capped at 40 pages; a page takes about 5 s on the shared gateway.
  • Billing records and charges are not extracted.

Eval results, nightly checks and cost per runVerify a run

How it's builtThe steps, the models and what each one checks
Self-host · your GPUs · recommended

Run it yourself, on request

  • The same open models and app, on 1x RTX PRO 6000 (96 GB) for Qwen3.8-27B, plus about 6 GB for the document reader and 10 GB for the embedder and reranker; the checks, merging and record run on CPU.
  • Data never leaves your machines, and there are no Decosa charges.
  • One prompt for Claude Code or Codex assembles the whole stack.
  • Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Hosted · by Decosa

Get an API key

  • Call the medical chronology with page cites API from your own code in minutes.
  • Every model answer carries a signed receipt.
  • Synthetic, public or test data only: real patient data belongs on your own hardware.
The open stack

A packet of medical records in, a dated chronology out: every line cites the file, page and box it came from.

For personal-injury and med-mal attorneys, life and disability insurers and APS reviewers. Give it the records as PDFs or images: typed notes, scanned and faxed forms, handwriting, lab tables. The document reader reads every page into elements with their boxes (and signs a receipt per file); the evidence retrieval block indexes each page's elements as chunks that keep page and box. Qwen3.8-27B reads each page, with the page image for scans, and lists encounters, diagnoses, procedures, imaging, labs, medications and referrals in a fixed schema. Code keeps an event only if its words are on the page and its date is written there (the dates block), turns ordered studies into referrals, merges copies of the same page from different providers, and a search per procedure finds other pages that give a different date; typed judgments settle the unclear pairs and say whether a condition documented before the injury is related to it. You get a Markdown chronology with conflicts, gaps in treatment and pre-existing conditions flagged, and a signed record of hashes. A paralegal and reviewer aid, not medical or legal advice.

Deployment
Self-host first
Regulatory
Written 27 Sep 2026. Medical records held by HIPAA covered entities (providers, health plans) and their business associates are protected health information; 45 CFR 160.103 names legal and consulting services among the services that make a vendor a business associate (https://www.law.cornell.edu/cfr/text/45/160.103, read 27 Sep 2026), and 45 CFR Part 164 Subpart E sets the privacy rules (https://www.hhs.gov/hipaa/for-professionals/privacy/index.html). A plaintiff's firm that gets records with the patient's authorization is usually not a business associate, and life and disability income insurance are excepted benefits outside HIPAA's 'health plan' (PHS Act 2791(c)(1), cited by 160.103; the list itself was not re-read today, unverified here), but state medical-privacy laws, protective orders and professional duties still apply to the same records. That is why real records are self-host only and the hosted demo refuses anything not marked synthetic. The tool lists what the records say with their page and box; it does not give medical opinions (causation, standard of care, prognosis) or legal advice, and a paralegal, nurse reviewer or attorney checks each line. Model licences: Apache-2.0 (Qwen3.8-27B, PaddleOCR-VL-1.6, Qwen3-Embedding, Qwen3-Reranker), MIT and Apache-2.0 (Docling and its layout weights).
Architecture
Text description

A record packet goes to decosa-api. The document reader reads every page into elements with page and box and signs a receipt per file. The retrieval block indexes each page's elements as chunks that keep page and box, with index and layout hashes, and copied pages are found. Qwen3.8-27B reads each page, with the image for scans, and lists typed events. Code keeps events whose quote is on the page and whose date is written there, turns orders into referrals, merges copies and mentions, and searches for other pages that give another date; typed judgments settle unclear pairs. The output is a dated chronology citing file, page and box, with conflicts, gaps and pre-existing conditions flagged, and a signed record of hashes.

Architecture

At a glance

Data retention
Nothing stored: files and page images live in memory for the request (the sample packet's reader output is cached in memory by file hash). The signed record holds hashes and receipt ids, never record text or the patient's name; logs carry counts only.
What leaves the box
Self-hosted on the direct route: nothing. The reader, the retrieval service, the model and the checks run on the same machine. Hosted: page text and page images go through the Decosa API, and only synthetic records are accepted.
What every line carries
Date, type, label, the quote, and a cite to file, page and box (plus the chunk and byte span in the index). Click-through to the box on the sample packet.
What it is not
Medical or legal advice. It does not judge causation, standard of care or damages; a paralegal, nurse reviewer or attorney checks each line at its box.
Quality tiers

Pick the tier for the quality you need

Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.

  • Lite

    document reader and the model, no retrieval service

    DECOSA_CHRONO_RETRIEVAL=bm25: the same reading, extraction, checks, merging and record; cites are located in a keyword index and the conflict search uses BM25 without the reranker, so about 10 GB less GPU memory.

    Models
    • decosa-api medical chronology (decosa_api/verticals/chronology), importing the document reader, evidence retrieval (scans into retrieval), typed judgment and dates blocks
    • Qwen3.8-27B (NVFP4)
    • Docling 2.130 with the Heron layout model (document reader block)
    • PaddleOCR-VL-1.6 (0.9B, document reader block)
    Hardware
    1x GPU with about 26 GB free for Qwen3.8-27B and the reader
    Quality evidence
    • Planted events found, conflicts, gaps with BM25 onlynot measured yetthe eval runner supports it (scripts/chronology_eval.py --bm25); not run on the test split
    Latency
    not measured separately; the reranked search is about a third of the merge step on the hosted run
    Verification
    Proof: strongSelf-host onlyEvery model call is receipted; the searches are BM25 in code, so there are no embed or rerank attestations.
  • In the hosted demo

    Standard

    reader, retrieval service and the model (hosted demo)

    The document reader, scans into retrieval with Qwen3-Embedding and Qwen3-Reranker, and Qwen3.8-27B per page with the page image. This is what the hosted demo runs, on synthetic records only.

    Models
    • decosa-api medical chronology (decosa_api/verticals/chronology), importing the document reader, evidence retrieval (scans into retrieval), typed judgment and dates blocks
    • Qwen3.8-27B (NVFP4)
    • Docling 2.130 with the Heron layout model (document reader block)
    • PaddleOCR-VL-1.6 (0.9B, document reader block)
    • Evidence retrieval block: Qwen3-Embedding-0.6B + Qwen3-Reranker-4B (decosa-retrieval service)
    Hardware
    1x RTX PRO 6000 96 GB (measured on shared cards)
    Quality evidence
    • Planted events found (held-out, 6 packets, 96 pages)189 / 198 (95.5%)decosa-api docs/evals/medical-chronology.md, test split, 27 Sep 2026
    • Date right / cite on the right page and box, of those founddates 189 / 189, cites 189 / 189decosa-api docs/evals/medical-chronology.md, test split
    • Lines that are not planted events2 / 224 (0.9%)decosa-api docs/evals/medical-chronology.md, test split
    • Planted date conflicts / gaps / pre-existing flagged6 / 6, 6 / 6, 12 / 12; 0 false flags of each kinddecosa-api docs/evals/medical-chronology.md, test split
    • Events on faxed copies carrying the copy's cite57 / 66decosa-api docs/evals/medical-chronology.md, test split
    Latency
    measured: several minutes per hundred pages on the shared gateway route (held-out eval); a couple of minutes for the sample.
    Verification
    Proof: strongEvery Qwen3.8 call is a separate gateway call with a gateway-signed receipt; each file has a signed document-reader receipt; each search has a signed search receipt with the index hash; the record lists them all.
Components

Every model in the stack

Models in this stack. Each row has a button that shows its licence, engine, verification and evidence.
ModelDetails
The event schema and page checks, copies, merging, conflicts, gaps, pre-existing, cites to chunks and bytes, the Markdown chronology and the signed record (no model; CPU)decosa-api medical chronology (decosa_api/verticals/chronology), importing the document reader, evidence retrieval (scans into retrieval), typed judgment and dates blocks
0 GBProof: partial
One call per page (the page's elements, plus the page image for a scan) listing typed events with quote and date; typed judgments for duplicate and conflict pairs and for pre-existing conditions; re-reads of regions the parser was unsure ofQwen3.8-27B (NVFP4)nvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab)
27.8B · 20 GBProof: strongIn the hosted demo
Finds and orders the regions of every page (text, headings, tables, checkboxes) with their boxesDocling 2.130 with the Heron layout model (document reader block)docling-project/docling-layout-heron on Hugging Face (opens in a new tab)
1 GBProof: partialIn the hosted demo
Reads each region of a scanned page (text, handwriting, tables as cells)PaddleOCR-VL-1.6 (0.9B, document reader block)PaddlePaddle/PaddleOCR-VL-1.6 on Hugging Face (opens in a new tab)
0.9B · 4.4 GBProof: partialIn the hosted demo
Scans into retrieval: each page's elements as chunks with page and box, the index and layout hashes, every cite located to a chunk and byte span, and a reranked search per procedure for other pages that give another dateEvidence retrieval block: Qwen3-Embedding-0.6B + Qwen3-Reranker-4B (decosa-retrieval service)Qwen/Qwen3-Reranker-4B on Hugging Face (opens in a new tab)
0.6B + 4B · 9 GBProof: partialIn the hosted demo

Around the models

Tools, services and hardware

Tools

Services

  • decosa-api:8445
    ${DECOSA_REGISTRY}/decosa-api:<tag>

    GET /chronology/info, /chronology/samples; POST /chronology/run (SSE or JSON); POST /retrieval/index-scans; POST /record/verify. Keeps no record text.

  • decosa-docreader (+ PaddleOCR-VL on vLLM):8497
    built from services/docreader; parser on vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1

    Page image in, regions with boxes out; nothing leaves the box.

  • decosa-retrieval:8499
    built from services/retrieval (no published image yet)

    Embedder and reranker for the evidence retrieval block, on the same GPU box.

  • vLLM (model):8114
    vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1

    Qwen3.8-27B NVFP4 behind our gateway (hosted) or called directly (self-host).

Hardware

  • 1x RTX PRO 6000 Blackwell 96 GB Fits

    Measured on our server: the model through the shared gateway on one card; the document reader (about 5.7 GB) and the retrieval service (about 10 GB) on another shared card.

  • 1x 48 GB card Fits

    Estimate: Qwen3.8-27B NVFP4 (about 20 GB with a small KV cache) plus about 16 GB for the reader and retrieval models. Not measured.

  • CPU only Does not fit

    The model and the page parser need a GPU. The checks, merging and record run on CPU.

Latency per lane

  • the 16-page sample packet, hosted gateway route, streamed144.4 s

    Measuredmeasured on our server 2026-09-27 (pre-release server): reading 51.6 s, extraction 43.0 s, merging, judgments and the conflict search 49.8 s; gateway shared with other workloads

  • per 100 pages, held-out eval (6 packets of 16 pages, one at a time)537.7 s

    Measuredmeasured on our server 2026-09-27: 537.7 s per 100 pages end to end (64 to 141 s a packet), shared gateway, 4 page calls at a time

  • smoke test: one 4-page file25.3 s

    Measuredmeasured on our server 2026-09-27, hosted gateway route

Notes

  • On six held-out synthetic packets (96 pages: typed, scanned, faxed and handwritten) it found 189 of 198 planted events; every found event had the right date and a cite on the right page and box.
  • It flagged all 6 planted date conflicts, all 6 gaps in treatment and all 12 conditions documented before the injury (related or not said right 12 of 12), with no false conflict, gap or pre-existing flag; 2 of 224 lines were not planted events.
  • Misses come mostly from the reader skipping a line on a scan (an ED diagnosis, an X-ray on a fax) and from the model leaving out the visit itself on a page.
  • The packets were written by the same author as the prompts, in one template family; real records are messier and have not been measured.
Assemble it

Run this exact stack on your machine

Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.

medical-chronology/assemble-prompt.md229 lines
# Assemble Decosa medical chronology with page cites on this machine

You are setting up a self-hosted medical chronology for a law firm, an insurer's claims or underwriting team, or a
record-review vendor. It reads a packet of medical records (typed PDFs, scans and faxes, handwriting, lab tables) and
returns:
- a dated chronology of encounters, diagnoses, procedures, imaging, labs, medications and referrals, every line citing
  the file, page and box it came from (and the chunk and byte span in the retrieval index);
- copies from different providers merged; conflicting dates for the same procedure, gaps in treatment after the date of
  injury, and conditions documented before it flagged;
- a signed, hash-chained record of hashes: each file and its document-reader receipt, the index and layout hashes, every
  entry's cites and quote hashes, the judgments and every model receipt id. No record text.

Work step by step, show me each command before running anything that needs sudo, and stop if a check fails.

**Before anything else, remind me:**
- Medical records are protected health information (HIPAA). On this box nothing leaves the machine: the document reader,
  the retrieval models, the language model and the checks all run here. Never point it at a hosted gateway while it
- This is a paralegal and reviewer aid, not medical or legal advice. A person checks each line at its box; it does not
  judge causation, standard of care or damages.

Repeat both points in your final summary.

## Step 0: set up with a coding agent, rehearse on mock data, then go private

This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:

1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
   nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
   "test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
   https://decosa.ai/samples/medical-chronology.zip (943 KB, 12 checks, synthetic or openly licensed: see `licence` in expected.json),
   show me what is in it, and run the rehearsal against the local API:
   `docker compose exec api python scripts/rehearse.py medical-chronology` (the api image carries the same bundle under /app/rehearsal/medical-chronology/;
   with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
   `python scripts/rehearse.py medical-chronology --bundle medical-chronology.zip --base-url http://127.0.0.1:<PORT>`.
   It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "every page was read and gave an answer", "the rotator cuff repair is listed and its conflicting date is flagged", "the 138-day gap in treatment is found"). Show me
   the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
   to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
   machine.

For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.

## What you are building

| service | image | model | port |
|---|---|---|---|
| `llm` | `${DECOSA_REGISTRY}/decosa-llm:0.1.0` (vLLM 0.29.0, `vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1`) | `nvidia/Qwen3.8-27B-NVFP4` @ `482ca0f3832238542f8f5295dde86b5f22711d80`, Apache-2.0 | internal 8000 |
| `parser` | the same vLLM image | `PaddlePaddle/PaddleOCR-VL-1.6` @ `c5630abae1d940eafe0697512a0325494b02ab42`, Apache-2.0 | internal 8498 |
| `docreader` | built from `decosa-api/services/docreader` | Docling layout `docling-project/docling-layout-heron` @ `8f39ad3c0b4c58e9c2d2c84a38465abf757272d8` (MIT + Apache-2.0) | internal 8497 |
| `retrieval` | built from `decosa-api/services/retrieval` | `Qwen/Qwen3-Embedding-0.6B` @ `97b0c614be4d77ee51c0cef4e5f07c00f9eb65b3` + `Qwen/Qwen3-Reranker-4B` @ `22e683669bc0f0bd69640a1354a6d0aebcfeede5`, Apache-2.0 | internal 8499 |
| `api` | `${DECOSA_REGISTRY}/decosa-api:0.1.0` (no GPU) | none | `127.0.0.1:8445` |

## 1. Check the GPU, driver and Docker

1. Run `nvidia-smi`. I need one NVIDIA GPU with at least 48 GB: the model about 20 GB plus KV cache, the page parser
   and layout model about 6 GB, the retrieval models about 10 GB. Driver 580 or newer.
   - Blackwell (RTX PRO 6000, B200): the defaults below (NVFP4) are the measured setup.
   - Hopper (H100/H200) or a 48 GB card: `LLM_MODEL=Qwen/Qwen3.8-27B-FP8`, `LLM_REVISION=main`, `LLM_GPU_UTIL=0.62`;
     on 48 GB also `LLM_MAX_LEN=32768` and `PARSER_GPU_UTIL=0.10`. Not measured.
   - Under 48 GB: tell me, and use the BM25 option in step 3 (no `retrieval` service).
2. Check `docker --version`, `docker compose version` and `docker run --rm --gpus all ubuntu nvidia-smi`. If Docker or
   the NVIDIA Container Toolkit is missing, install them from the official Docker and NVIDIA repositories, run
   `sudo nvidia-ctk runtime configure --runtime=docker` and restart Docker.
3. Confirm about 90 GB of free disk.

## 2. Get the images and the source

The images are **on request** while self-host is in early access: ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry (set it as `DECOSA_REGISTRY`), pull access and the compose file. Try `docker pull ${DECOSA_REGISTRY}/decosa-{llm,api}:0.1.0`. If a pull
fails, clone `decosa-api` into `~/decosa-chrono/decosa-api` and build the api with
`docker build -f docker/api/Dockerfile -t ${DECOSA_REGISTRY}/decosa-api:0.1.0 .`. The `docreader` and `retrieval`
services are always built from the clone, so clone it either way. If neither works, stop and tell me.

## 3. Write the compose file

Create `~/decosa-chrono/.env`:

```bash
DECOSA_TAG=0.1.0
DECOSA_GPU=0
LLM_MODEL=nvidia/Qwen3.8-27B-NVFP4
LLM_REVISION=482ca0f3832238542f8f5295dde86b5f22711d80
LLM_MAX_LEN=65536
LLM_GPU_UTIL=0.62
PARSER_GPU_UTIL=0.05
DECOSA_SIGNER_NAME="<who signs these records, e.g. Example LLP records team>"
```

Create `~/decosa-chrono/docker-compose.yml` with exactly these services:

```yaml
name: decosa-chrono
x-gpu: &gpu { resources: { reservations: { devices: [ { driver: nvidia, device_ids: ["${DECOSA_GPU:-0}"], capabilities: [gpu] } ] } } }
x-health: &health { interval: 15s, timeout: 5s, retries: 5 }
services:
  llm:
    image: ${DECOSA_REGISTRY}/decosa-llm:${DECOSA_TAG}
    deploy: *gpu
    ipc: host
    restart: unless-stopped
    volumes: [hf-cache:/root/.cache/huggingface]
    command: ["${LLM_MODEL}", "--revision", "${LLM_REVISION}", "--served-model-name", "qwen3.8-27b",
              "--max-model-len", "${LLM_MAX_LEN}", "--gpu-memory-utilization", "${LLM_GPU_UTIL}", "--max-num-seqs", "16",
              "--kv-cache-dtype", "fp8_e4m3", "--speculative-config", '{"method":"mtp","num_speculative_tokens":3}',
              "--seed", "0", "--enable-force-include-usage", "--host", "0.0.0.0", "--port", "8000"]
    healthcheck: { <<: *health, test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8000/health', timeout=4)"], start_period: 900s }
  parser:
    image: vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1
    deploy: *gpu
    ipc: host
    restart: unless-stopped
    volumes: [hf-cache:/root/.cache/huggingface]
    command: ["PaddlePaddle/PaddleOCR-VL-1.6", "--revision", "c5630abae1d940eafe0697512a0325494b02ab42", "--trust-remote-code",
              "--max-model-len", "8192", "--gpu-memory-utilization", "${PARSER_GPU_UTIL}", "--max-num-seqs", "16",
              "--no-enable-prefix-caching", "--mm-processor-cache-gb", "0", "--generation-config", "vllm", "--host", "0.0.0.0", "--port", "8498"]
    healthcheck: { <<: *health, test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8498/health', timeout=4)"], start_period: 600s }
  docreader:
    build: { context: ./decosa-api/services/docreader }
    deploy: *gpu
    restart: unless-stopped
    depends_on: { parser: { condition: service_healthy } }
    environment:
      DOCREADER_PARSER_URL: http://parser:8498/v1
      DOCREADER_PARSER_MODEL: PaddlePaddle/PaddleOCR-VL-1.6
      DOCLING_DEVICE: cuda
    volumes: [hf-cache:/models]
  retrieval:
    build:
      context: ./decosa-api/services/retrieval
      dockerfile_inline: |
        FROM python:3.12-slim
        WORKDIR /app
        COPY requirements.txt .
        RUN pip install --no-cache-dir -r requirements.txt --extra-index-url https://download.pytorch.org/whl/cu130
        COPY *.py .
        CMD ["python", "server.py"]
    deploy: *gpu
    restart: unless-stopped
    environment: { RETRIEVAL_HOST: 0.0.0.0, RETRIEVAL_PORT: "8499", RETRIEVAL_EMBED: qwen3-emb-0.6b, RETRIEVAL_RERANK: qwen3-rr-4b, HF_HOME: /root/.cache/huggingface }
    volumes: [hf-cache:/root/.cache/huggingface]
    healthcheck: { <<: *health, test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8499/health', timeout=4)"], start_period: 600s }
  api:
    image: ${DECOSA_REGISTRY}/decosa-api:${DECOSA_TAG}
    restart: unless-stopped
    depends_on: { llm: { condition: service_healthy }, docreader: { condition: service_healthy }, retrieval: { condition: service_healthy } }
    environment:
      DECOSA_LLM_ROUTE: direct                  # local model only; receipts are signed by this box's key ("attested")
      DECOSA_LLM_URL: http://llm:8000/v1
      DECOSA_LLM_MODEL: qwen3.8-27b
      DECOSA_LOCAL_SIGNING: "on"                # Ed25519 key created at /data/attest/ed25519.pem on first start
      DECOSA_SIGNER_NAME: ${DECOSA_SIGNER_NAME}
      DECOSA_DOCREADER_URL: http://docreader:8497
      DECOSA_RETRIEVAL_URL: http://retrieval:8499
      DECOSA_CHRONO_SYNTHETIC_ONLY: "0"         # this box accepts real records
      DECOSA_DOCREADER_SYNTHETIC_ONLY: "0"
      DECOSA_CHRONO_MAX_PAGES: "600"            # pages per run (hosted demo: 40)
      DECOSA_SESSIONS_PER_IP_HOUR: "1000"
      DECOSA_BUDGET_LLM_TOKENS: "600000"        # per session; about 350 generated tokens a page
      DECOSA_SESSION_TTL_S: "28800"
      DECOSA_LLM_TIMEOUT_S: "300"
      DECOSA_CORS_ORIGIN_REGEX: '^https?://(localhost|127\.0\.0\.1)(:\d+)?$$'
    ports: ["127.0.0.1:8445:8445"]
    volumes: [decosa-data:/data]
    healthcheck: { <<: *health, test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8445/healthz', timeout=4)"], start_period: 20s }
volumes: { hf-cache: {}, decosa-data: {} }
```

**No room for the retrieval models?** Delete `retrieval` and its `depends_on` entry, drop `DECOSA_RETRIEVAL_URL`, and add
`DECOSA_CHRONO_RETRIEVAL: bm25` to the api: cites are still located in the index by keyword chunks; the conflict sweep
uses BM25 instead of the reranker.

Use the named volume `decosa-data` exactly as written: a root-owned host bind mount makes the API fail on
`/data/keys.sqlite`. Run `docker compose up -d --build` and poll `docker compose ps` until every service is healthy (the
LLM takes 5 to 10 minutes the first time). `curl -s localhost:8445/chronology/info | jq '{synthetic_only, blocks}'` should
show `synthetic_only: false`, the document reader reachable and retrieval `hybrid+rerank`.

## 4. Smoke test on the bundled synthetic packet

```bash
API=localhost:8445
TOKEN=$(curl -s $API/demo/session -H 'content-type: application/json' -d '{"vertical":"medical-chronology"}' | jq -r .token)
curl -s $API/chronology/run -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' -d '{"sample":"whitlock"}' > /tmp/chrono.json
jq -r '.entries[] | "\(.date)\t\(.type)\t\(.label[0:50])\t\(.cites[0].cite)\t\(.flags|join(","))"' /tmp/chrono.json
jq '{counts, gaps: [.gaps[] | {from, to, days}], preexisting: [.preexisting[] | {date, label, related}], copies}' /tmp/chrono.json
jq '{record}' /tmp/chrono.json | curl -s $API/record/verify -H 'content-type: application/json' -d @- | jq '{ok, summary}'
```

Pass if: the procedure "right shoulder arthroscopic rotator cuff repair" on 2024-04-28 carries the flag `conflict` with
the other date 2024-05-05; there is one gap from 2024-05-12 to 2024-09-27; the 2022-11-15 low back pain diagnosis is
flagged before the injury and judged related; the pages F1 p4 and F5 p2 are listed as copies; every entry has at least one
cite with a box; and the record verifies. About 20 to 40 entries is normal; the planted answer key has 32 required events.

## 5. Run a real packet

```bash
python3 - <<'PY' > packet.json
import base64, json, pathlib
files = [{"name": p.name, "file_b64": base64.b64encode(p.read_bytes()).decode()} for p in sorted(pathlib.Path("records").iterdir())
         if p.suffix.lower() in (".pdf", ".png", ".jpg", ".jpeg", ".tif", ".tiff")]
print(json.dumps({"files": files, "incident_date": "2024-03-08", "claimed": ["low back", "right shoulder"], "gap_days": 60}))
PY
curl -s $API/chronology/run -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' -d @packet.json > run.json
jq -r .markdown run.json > chronology.md
jq .record run.json > chronology-record.json
```

Limits per run: 12 files of 12 MB, and `DECOSA_CHRONO_MAX_PAGES` pages in all. Measured on our server's shared gateway:
about 5 to 7 s a page end to end; a private GPU runs faster. For a very large packet, split it by provider and run
each, or raise the limits. Check `complete` and `incomplete_pages` in every result.

## 6. Point tools at the local API

- Base URL `http://localhost:8445`; for a web app set `NEXT_PUBLIC_DECOSA_API=http://localhost:8445`.
- `POST /chronology/run` returns JSON, or Server-Sent Events with `Accept: text/event-stream`. `POST /docreader/read` and
  `POST /retrieval/index-scans` are the two building blocks on their own.
- Keep the chronology and the record with the case file. Anyone can re-check the record with `POST /record/verify`
  against the key at `GET /attest/signing-key`.
- Keep the API on 127.0.0.1; for other users, put a TLS reverse proxy with authentication in front, and handle the box
  under your HIPAA security program (access control, audit logs, encryption at rest).

## 7. Keep the direct route

`DECOSA_LLM_ROUTE=gateway` would send page text and images to the hosted Decosa
API. Leave it off: this box holds PHI.

Finish with a summary: what is running, the health output, the smoke-test results, and the reminders above.
Rules and regulations it checks againstDated, linked to the primary source; not legal advice

Regulation watch

Loading the watch status…

4 laws, rules and guidance pages cited; 3 watched nightly at the primary source. A change marks this page for a human re-check; nothing is edited automatically. What we cite and how it is watched

Technical detailsModels, where it runs, labels

In short

Last reviewed

What it is
A medical chronology tool for record reviewers: every line of the chronology cites the file, page and box it came from, and copies, conflicting dates, gaps in treatment and pre-existing conditions are flagged.
Who it's for
Personal-injury and med-mal attorneys and paralegals, nurse reviewers, and life and disability insurers reviewing attending physician statements.
Where it runs
Self-host for real records (PHI); the hosted demo takes synthetic records only
Key numbers

On six held-out synthetic packets (96 pages) it found 189 of 198 planted events, each with the right date and a cite on the right page and box, and 2 of 224 lines were not planted events; real records are not measured yet.

  • 189 / 198 Planted events found (test split, n = 198)
  • 189 / 189 Date right, of events found (test split, n = 189)
  • 189 / 189 Cite on the right page and box, of events found (test split, n = 189)
  • 25.3 s Median end-to-end run, hosted (QA sweep 2026-09-27)
All results, datasets and caveats
Models
Qwen3.8-27B (events per page with the page image, typed judgments) · document reader (Docling layout + PaddleOCR-VL) · Qwen3-Embedding + Qwen3-Reranker (retrieval over the page boxes)
Where
Self-host for real records (PHI); the hosted demo takes synthetic records only
Checks
Receipt per model call; a signed document-reader receipt per file; retrieval index and layout hashes with signed search receipts; every quote found on its page; signed hash-chained record
Industry
Legal · Healthcare
Output
Structured data · Signed record or verdict
Data
Patient data (PHI)
Hardware
1× 96 GB GPU
Licence
Permissive (Apache-2.0, MIT)
Runs in
Self-host

Questions people ask

Can it build a medical chronology from scanned and handwritten records?

Yes: every page goes through the document reader, which reads scans, faxes, handwriting and tables and gives each element a box. On synthetic packets with scanned, faxed and handwritten pages it found 189 of 198 planted events. Handwriting is read less reliably than print, so check handwritten dates at the cited box.

Can it make up an event?

An event reaches the chronology only if its quoted words are found on the page it cites and its date is written there; everything else is dropped or flagged. On the held-out packets 2 of 224 lines were not planted events (therapy techniques from a plan listed as procedures).

Do patient records leave our network?

Not when you self-host: the document reader, the retrieval service and the model all run on your own GPU box. The hosted demo accepts synthetic records only and keeps nothing.

Does it find pre-existing conditions and gaps in treatment?

Give it the date of injury and the claimed injuries. It flags diagnoses documented before that date, with a typed judgment on whether each is related, and any gap between visits longer than the gap you set (60 days by default). On the held-out packets it found 12 of 12 pre-existing conditions and 6 of 6 gaps with no false flags.

Is this medical or legal advice?

No. It is a paralegal and reviewer aid: it lists what the records say with the page and box, and a person checks each line. It does not judge causation, standard of care or damages.

Ask a question or leave feedbackWe read every message and publish useful answers
Questions & feedback

Ask about Medical chronology with page cites

We read every message. Questions, comments and our answers show here once we have reviewed and approved them.

Loading questions…

This is a

Plain text. Please leave out personal, patient or client data.

Shown with your message if we publish it. Leave blank to post as “A visitor”.

Nothing appears here until we have read and approved it.