Self-host
Run it on your own hardware
Same apps, same models, your hardware. Your data stays on your machine, and there are no Decosa usage charges. Self-host is on request while it's in early access: the code and images aren't public yet.
A few tools look up citations, drug names or DOIs in public services even when self-hosted; none sends your documents. Where your data goes, tool by tool
No NVIDIA GPU? Most text and voice tools run natively on an Apple Silicon Mac, with one script and no Docker. Run it on a Mac Studio
Hardware check: 93 tools
- CPU only, 64 GB RAM3 run24 smaller tier66 don't fit
- GeForce RTX 409058 run18 smaller tier17 don't fit
- GeForce RTX 509058 run17 smaller tier18 don't fit
- 2x GeForce RTX 509084 run8 smaller tier1 don't fit
- L40S67 run17 smaller tier9 don't fit
- H100 80 GB (SXM)88 run3 smaller tier2 don't fit
- RTX PRO 6000 Blackwell 96 GB91 run2 smaller tier0 don't fit
- 2x RTX PRO 6000 Blackwell 96 GB91 run2 smaller tier0 don't fit
- Apple M3 Ultra (Mac Studio), 96 GB72 run12 smaller tier0 don't fit9 can't tell
- Apple M5 Max, 64 GB71 run13 smaller tier0 don't fit9 can't tell
Pick your own hardware for a tier, model swaps and a setup prompt for each tool.
Running a therapy practice? The audit-ready practice workflow has its own guide: self-host the practice.
On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.
- 1
Check the GPU, Docker and the NVIDIA Container Toolkit
The driver must see the GPU, and Docker must be able to pass it into a container.
nvidia-smi docker compose version docker run --rm --gpus all ubuntu nvidia-smi
- 2
Fetch the compose file
One file describes the API, the speech model and the language model as services.
mkdir -p ~/decosa && cd ~/decosa curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml - 3
Pull and start
The first start downloads pinned model weights, tens of gigabytes.
docker compose pull docker compose up -d
- 4
Check health
Wait until the API reports ok with both models loaded. Then point your app at the local base URL.
curl -fsS http://localhost:<PORT>/healthz # {"ok": true, "asr": true, "llm": true, ...}
Set up with a coding agent, rehearse on mock data, then go private
- Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
- Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's
expected.json. Every check must print PASS. - Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
docker compose exec api python scripts/rehearse.py <use-case-id>
Each tool's Self-host tab links its own bundle: synthetic or openly licensed inputs and an expected.json of checkable results. The list is at /samples/index.json, and the same bundles ship inside the api image.
A decosa command-line installer (init, doctor, up) is planned. Docker Compose is the supported route for now.
Why patient data should stay on site
- Visit audio and notes are protected health information. Running the scribe on a GPU in the clinic means that audio and text never cross the internet to a vendor.
- The hosted demo keeps sessions in memory and deletes them when they end, but it is a public demo. It must not receive PHI.
- A local box keeps working when the internet link is down, and latency stays low.
- A self-hosted box signs its own receipts with a key it generates on first start: each note carries the model name, its weights root and the hashes of what went in and came out, without sending anything anywhere. That is an attestation by the clinic’s own box, not a proof that the model ran.
Pick a tool
Each tool page has a copy-paste prompt that has your coding agent install Docker and the NVIDIA Container Toolkit if needed, pull the containers, start them and check health, then rehearse on that tool's mock-data bundle before any real data goes near it.
- Run the visit with a copilot1× RTX PRO 6000 (96 GB), or 2× RTX 5090
- Run an open model behind the OpenAI API1× RTX PRO 6000 (96 GB), or 2× RTX 5090
- Make a story, a song or a music video1× RTX PRO 6000 (96 GB)
- Write the inspection report1× RTX PRO 6000 (96 GB), or 2× RTX 5090
- Translate a talk live1× RTX PRO 6000 (96 GB), or 2× RTX 5090
- Make a tamper-evident meeting record1× RTX PRO 6000 (96 GB), or 2× RTX 5090
- Check your provider serves the model you pay forNo GPU needed to audit. References are recorded on 1× RTX PRO 6000 (96 GB).
- Make a disclosed UGC adHosted renders on fal (no local GPU); self-host: 1× RTX PRO 6000 (96 GB)
- Catch AI answers your sources don't back1× RTX PRO 6000 (96 GB) or 1× RTX 5090 (32 GB) for the judge; the checker itself runs on CPU
- Digest a deposition1× RTX PRO 6000 (96 GB); no GPU for the parser, flags and export
- Check a brief before filing1× RTX PRO 6000 (96 GB) or 1× RTX 5090 (32 GB) for the judge; parsing, lookups, quotes and the privacy check run on CPU
- Pre-check promo claims for MLR1× RTX PRO 6000 (96 GB) for the model; extraction, rules and the packet run on CPU
- Check if an open model can take over your prompt1× RTX PRO 6000 (96 GB) or 1× RTX 5090 (32 GB) for the candidate and judge; scoring runs on CPU
- Classify with a confidence you can act on1× RTX PRO 6000 (96 GB) or 1× RTX 5090 (32 GB) for the model; the API runs on CPU
- Answer a security questionnaire1× RTX PRO 6000 (96 GB) or 1× RTX 5090 (32 GB) for the model; parsing, checks and export run on CPU
- Prove what your agent did1× RTX PRO 6000 (96 GB) for the decision model; the recorder and browser run on CPU
- Run an end-to-end browser test1× RTX PRO 6000 (96 GB) for the action model; the runner, browser and verifier run on CPU
- Build a model-risk evidence pack1× RTX PRO 6000 (96 GB) or 1× RTX 5090 (32 GB) for the grader; the pack runner is CPU; the system under test runs wherever it runs
- Check an AI note1× RTX PRO 6000 (96 GB) or 1× RTX 5090 (32 GB) for the judge; the diarizer (optional, for audio) needs a second GPU slot
- Build a privilege log1× RTX PRO 6000 (96 GB) or 1× RTX 5090 (32 GB) for the model; the people map, grouping and leak rules run on CPU
- Log human edits and editor sign-off1× RTX PRO 6000 (96 GB) or 1× RTX 5090 (32 GB) for the model; metrics, ledger and verification run on CPU
- Check a police report against bodycam1× RTX PRO 6000 (96 GB) for Qwen3.8-27B and the diarizer; the check itself needs no GPU
- Redline from firm precedents1× RTX PRO 6000 (96 GB) or 1× RTX 5090 (32 GB) for the model; the DOCX writer, search and checks run on CPU
- Redact a public-records release1× RTX PRO 6000 (96 GB) or 1× RTX 5090 (32 GB) for the model; pattern finders, the release PDF and its check run on CPU
- Score an oral exam against a rubric1× RTX PRO 6000 (96 GB) for Voxtral, the diarizer and Qwen3.8-27B; text-only scoring needs only the Qwen server
- Chart claim support in a patent spec1× RTX PRO 6000 (96 GB) or 1× RTX 5090 (32 GB) for the model; parsing and the code checks run on CPU
- Make a cleared songAbout 50 GB free on one 96 GB card for MiniMax-Music3 (ACE-Step needs about 15 GB); the similarity check runs on CPU
- Pre-check samples and lyricsAny CPU for the audio matching and lyric spans; Qwen3.8-27B (1× RTX 5090 32 GB or larger) for the lyric labels
- Check split sheets and metadata1× RTX PRO 6000 (96 GB) or 1× RTX 5090 (32 GB) for the model; parsing, OCR and every check run on CPU
- Check a paper's citations1× RTX PRO 6000 (96 GB) or 1× RTX 5090 (32 GB) for the model; parsing and the reference checks run on CPU
- Keep a signed lab notebook1× RTX PRO 6000 (96 GB) or 1× RTX 5090 (32 GB) for the analysis model; the notebook, signatures and verification run on CPU
- Check green claims in copy1× RTX PRO 6000 (96 GB) or 1× RTX 5090 (32 GB) for the model; the rulepack, grounding spans and the report run on CPU
- Record F&I add-on disclosures1× RTX PRO 6000 (96 GB) for Qwen3.8-27B and the diarizer; the jacket rules and the record need no GPU
- Dub a video in your own voice1× RTX PRO 6000 (96 GB) for Qwen3.8-27B and Voxtral; the voice model needs about 4 GB of GPU or runs on CPU at about 4× real time
- Pre-flight a synthetic-performer adAny CPU for the video steps (ffmpeg, Tesseract, C2PA); Qwen3.8-27B (1× RTX 5090 32 GB or larger) for the script checks
- Turn a script into an animatic1× RTX PRO 6000 (96 GB): Qwen3.8-27B for the shot list, and about 34 GB for Wan2.2-VACE-Fun-A14B (fp8) in ComfyUI; the voices run on CPU
- Produce an audio dramaQwen3.8-27B (1× RTX 5090 32 GB or larger) for the parse; the voices, sound and mix run on CPU; ACE-Step only if you compose new music (a GPU with about 10 GB free)
- Make a music video for your track1× RTX PRO 6000 (96 GB) shared with the text model, or a 48 GB card for the video model plus a remote text model; the analysis runs on CPU
- Write a SAR narrative1× RTX PRO 6000 (96 GB) or 1× RTX 5090 (32 GB) for Qwen3.8-27B; the numeric check, the screen and the report run on CPU
- Review a claim file for conduct1x RTX PRO 6000 (96 GB) for Qwen3.8-27B; the deadline rules and the record need no GPU
- Review a collections call1× RTX PRO 6000 (96 GB) for Qwen3.8-27B and the diarizer; the call-log maths and the record need no GPU
- Draft breach notices and deadlines1× RTX PRO 6000 (96 GB) or 1× RTX 5090 (32 GB) for Qwen3.8-27B; the clocks, the time and number checks and the record run on CPU
- Tie out MD&A figuresAny CPU for the tie-out; 1× RTX PRO 6000 (96 GB) or 1× RTX 5090 (32 GB) for Qwen3.8-27B if you want claims in words read
- Disposition a sanctions alert1× RTX PRO 6000 (96 GB) or 1× RTX 5090 (32 GB) for Qwen3.8-27B; the comparison, the proposal and the record run on CPU
- Review a Medicare sales call1× RTX PRO 6000 (96 GB) for Qwen3.8-27B and the diarizer; the call-sheet items and the record need no GPU
- Appeal a denial1x RTX PRO 6000 (96 GB) for Qwen3.8-27B; the deadline rules and the signed packet need no GPU
- Check if a video is made for kidsQwen3.8-27B (1× RTX 5090 32 GB or larger); speech to text for videos needs MOSS-Transcribe-Diarize on a GPU; OCR and the record run on CPU
- Check HCC codes for MEAT evidence1× RTX PRO 6000 (96 GB) or 1× RTX 5090 (32 GB) for Qwen3.8-27B; the V28 map, the record checks and the signed record run on CPU
- Draft VEX for scanner findings1x RTX PRO 6000 (96 GB) for Qwen3.8-27B; the collector, the checks, rules mode and the signing run on CPU
- Triage a device complaint for MDR1× RTX PRO 6000 (96 GB) for Qwen3.8-27B; the clocks, the file check and the signed records run on CPU
- Map evidence to NIST 800-1711× RTX PRO 6000 (96 GB) or 1× RTX 5090 (32 GB) for Qwen3.8-27B; the catalog, artefact checks, POA&M rules and signed record run on CPU
- Check a lay summary's numbersQwen3.8-27B (1× RTX 5090 32 GB or larger); the number tracing, checks and record run on CPU
- Investigate a Reg E dispute1x RTX PRO 6000 (96 GB) for Qwen3.8-27B; the clocks, the checklist and the signed record need no GPU
- Stage a listing photo with disclosureWan2.2-VACE-Fun-A14B (fp8) and Qwen3.8-27B: one 96 GB GPU (measured on two shared ones); the check, label, credential and record run on CPU
- Write a due-diligence red-flag memo1x RTX PRO 6000 (96 GB) for Qwen3.8-27B plus about 10 GB for the embedder and reranker; the quote checks, memo and record run on CPU
- Draft a tariff classification memo1x RTX PRO 6000 (96 GB) for Qwen3.8-27B; the embedder and reranker fit in about 10 GB beside it or on a second card; the checks and signing run on CPU
- Check CSR numbers against tables1x RTX PRO 6000 (96 GB) for Qwen3.8-27B, plus about 10 GB for the embedder and reranker and about 6 GB for the document reader; the comparisons, report and record run on CPU
- Build an EU GPSR listing1x RTX PRO 6000 (96 GB): Qwen3.8-27B NVFP4 plus Hy-MT2-7B (about 18 GB); the checks run on CPU
- Build a medical chronology1x RTX PRO 6000 (96 GB) for Qwen3.8-27B, plus about 6 GB for the document reader and 10 GB for the embedder and reranker; the checks, merging and record run on CPU
- Screen a claim for Federal IDR1x RTX PRO 6000 (96 GB) for Qwen3.8-27B; the clocks, the screen from structured facts and the signed packet need no GPU
- Quote a job from a walkthrough videoQwen3.8-27B with video input and MOSS-Transcribe-Diarize: one 96 GB GPU (measured on two shared ones); ffmpeg, pricing and the record on CPU
- Turn an expert video into an SOPQwen3.8-27B with video input and MOSS-Transcribe-Diarize: one 96 GB GPU (measured on two shared ones); ffmpeg, keyframes and the record on CPU
- Check generated product images1x RTX PRO 6000 (96 GB) for Qwen3.8-27B with its vision tower; the document reader's layout model and parser run beside it; the checks, credential and record on CPU
- Draft an adverse-event case1x RTX PRO 6000 (96 GB) for Qwen3.8-27B, Hy-MT2-7B, the call recogniser and the document reader; the criteria, clocks and records run on CPU
- Compare safety sections across labels1x RTX PRO 6000 (96 GB): Qwen3.8-27B NVFP4 plus Hy-MT2-7B (about 18 GB) and the document reader (about 6 GB); sections, number checks and the record run on CPU
- Find accessibility fixes for a shop1x RTX PRO 6000 (96 GB) for Qwen3.8-27B with its vision tower; headless Chromium and axe-core on CPU
- Fill a form from your own papers1x RTX PRO 6000 (96 GB) for Qwen3.8-27B; the sandboxed Chromium, form reader and network hold run on CPU
- Turn a client call into notes1x RTX PRO 6000 (96 GB) for the language model; Voxtral and the diarizer need about 27 GB more (a second card on our server)
- Check the other side's brief1x RTX PRO 6000 (96 GB) or 1x RTX 5090 (32 GB) for the judge; the hidden-text scan, lookups and quotation match run on CPU
- Check their discovery responses1x RTX PRO 6000 (96 GB) for Qwen3.8-27B; splitting, rule cites, deadlines, the letter and the record run on CPU
- Answer a payer audit1x RTX PRO 6000 (96 GB) for Qwen3.8-27B; the deadline, the matching, the packet and the signed record need no GPU
- Check a prior auth before you send it1x RTX PRO 6000 (96 GB) for Qwen3.8-27B; the policy reader, the decision clock and the signed record need no GPU
- Write a note from my jottings1x RTX PRO 6000 (96 GB) or 1x RTX 5090 (32 GB, estimate) for Qwen3.8-27B; the photo's second reader needs about 3 GB more; typed jottings need nothing else
- Capture audit evidence from your admin screensThe quarterly run needs only a CPU (headless Chromium). Setup and repairs call Qwen3.8-27B: 1x RTX PRO 6000 (96 GB) self-hosted, or the hosted gateway
- Put the visit into your EHR as drafts1x RTX PRO 6000 (96 GB) for Qwen3.8-27B; the medicine check and the commit detector run on CPU
- Read a demand before the deadline1x RTX PRO 6000 (96 GB) for Qwen3.8-27B; the deadline rules and the tie-out need no GPU
- Check a certificate request1x RTX PRO 6000 (96 GB) for Qwen3.8-27B
- Check a bank-detail change before you pay1x RTX PRO 6000 (96 GB) for Qwen3.8-27B; the header and vendor-file checks run on CPU
- Turn a family interview into a filmCPU for the edit; speech recognition, translation and the photo-back reader share one GPU (no video model)
- Make a music video starring you1x 96 GB card for MiniMax H3 shots (fp8, about 50 GB); the storyboard runs on FLUX.2 klein (about 6 GB) and the edit on CPU
- Make a film of your story together1x 96 GB card for MiniMax H3 shots; the storyboard runs on FLUX.2 klein and the edit on CPU
- Make a settlement video1x RTX PRO 6000 (96 GB) for Qwen3.8-27B and the document reader (as for the medical chronology); the voices, the checks and the renderer run on CPU
- Find the themes in your interviews1x RTX PRO 6000 (96 GB) for Qwen3.8-27B and the diarizer; the voice check, the quote check and your own model run on CPU
- Find the songs in your mixAny CPU; no GPU
- Reply to a review without breaking patient privacy1x RTX PRO 6000 (96 GB) for Qwen3.8-27B; the checks run on CPU
- See what studies found for a supplement1x RTX PRO 6000 (96 GB) or 1x RTX 5090 (32 GB) for Qwen3.8-27B; about 8 GB for the reranker; the search, checks and grade run on CPU