Skip to content
decosa

Self-host

Run it on your own hardware

Same apps, same models, your hardware. Your data stays on your machine, and there are no Decosa usage charges. Self-host is on request while it's in early access: the code and images aren't public yet.

A few tools look up citations, drug names or DOIs in public services even when self-hosted; none sends your documents. Where your data goes, tool by tool

No NVIDIA GPU? Most text and voice tools run natively on an Apple Silicon Mac, with one script and no Docker. Run it on a Mac Studio

Hardware check: 93 tools

  • CPU only, 64 GB RAM3 run24 smaller tier66 don't fit
  • GeForce RTX 409058 run18 smaller tier17 don't fit
  • GeForce RTX 509058 run17 smaller tier18 don't fit
  • 2x GeForce RTX 509084 run8 smaller tier1 don't fit
  • L40S67 run17 smaller tier9 don't fit
  • H100 80 GB (SXM)88 run3 smaller tier2 don't fit
  • RTX PRO 6000 Blackwell 96 GB91 run2 smaller tier0 don't fit
  • 2x RTX PRO 6000 Blackwell 96 GB91 run2 smaller tier0 don't fit
  • Apple M3 Ultra (Mac Studio), 96 GB72 run12 smaller tier0 don't fit9 can't tell
  • Apple M5 Max, 64 GB71 run13 smaller tier0 don't fit9 can't tell

Pick your own hardware for a tier, model swaps and a setup prompt for each tool.

Running a therapy practice? The audit-ready practice workflow has its own guide: self-host the practice.

On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.

  1. 1

    Check the GPU, Docker and the NVIDIA Container Toolkit

    The driver must see the GPU, and Docker must be able to pass it into a container.

    nvidia-smi
    docker compose version
    docker run --rm --gpus all ubuntu nvidia-smi
  2. 2

    Fetch the compose file

    One file describes the API, the speech model and the language model as services.

    mkdir -p ~/decosa && cd ~/decosa
    curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml
  3. 3

    Pull and start

    The first start downloads pinned model weights, tens of gigabytes.

    docker compose pull
    docker compose up -d
  4. 4

    Check health

    Wait until the API reports ok with both models loaded. Then point your app at the local base URL.

    curl -fsS http://localhost:<PORT>/healthz
    # {"ok": true, "asr": true, "llm": true, ...}

Set up with a coding agent, rehearse on mock data, then go private

  1. Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
  2. Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's expected.json. Every check must print PASS.
  3. Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
Rehearsal command
docker compose exec api python scripts/rehearse.py <use-case-id>

Each tool's Self-host tab links its own bundle: synthetic or openly licensed inputs and an expected.json of checkable results. The list is at /samples/index.json, and the same bundles ship inside the api image.

A decosa command-line installer (init, doctor, up) is planned. Docker Compose is the supported route for now.

Why patient data should stay on site

  • Visit audio and notes are protected health information. Running the scribe on a GPU in the clinic means that audio and text never cross the internet to a vendor.
  • The hosted demo keeps sessions in memory and deletes them when they end, but it is a public demo. It must not receive PHI.
  • A local box keeps working when the internet link is down, and latency stays low.
  • A self-hosted box signs its own receipts with a key it generates on first start: each note carries the model name, its weights root and the hashes of what went in and came out, without sending anything anywhere. That is an attestation by the clinic’s own box, not a proof that the model ran.

Pick a tool

Each tool page has a copy-paste prompt that has your coding agent install Docker and the NVIDIA Container Toolkit if needed, pull the containers, start them and check health, then rehearse on that tool's mock-data bundle before any real data goes near it.