Reply to a review without breaking patient privacy
A short reply to read and post, checked in code for offers, invented facts and anything that confirms a patient, with the blocked words shown.
Built on: Site checks, Signed record
Loading the tool…
Use it your way
Use it from your codeThe hosted API with your key, and prompts to paste into a coding agent
Get an API key
- Call the review reply with patient privacy API from your own code in minutes.
- Every model answer carries a signed receipt.
- Nothing to install; we run the models.
Run it yourself, on request
- The same open models and app, on 1x RTX PRO 6000 (96 GB) for Qwen3.8-27B; the checks run on CPU.
- Data never leaves your machines, and there are no Decosa charges.
- One prompt for Claude Code or Codex assembles the whole stack.
- Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
Build with it
Paste one of these into Claude Code, Codex or another coding agent. The first wires your project to the hosted API with your DECOSA_API_KEY. The second pulls our containers and runs the same stack on your own GPU, with no Decosa charges.
- Base URL
- https://api.decosa.ai
- Auth
Authorization: Bearer $DECOSA_API_KEY(or a demo session token)- Tool id
- review-reply
Use the hosted API
# Decosa Reply to a review without breaking patient privacy: use the hosted API
You are wiring Decosa's review-reply tool into this project. It returns:
- a short draft reply to a public review, ready for the owner to read and edit;
- the checks it passed, in code: no offers, no contact details or numbers the owner didn't give, no placeholders, not defensive, negative reviews taken offline;
- for a health or care business, the patient-privacy rules: nothing that confirms the reviewer is a patient or mentions a visit, treatment, condition, wait, bill, family member or name, plus a second check of the reply by the open model;
- the words it blocked, and a signed record that keeps only the review's fingerprint (SHA-256).
Every model call has a signed receipt. Use only what is listed below; if you need something else, stop and ask me.
- Base URL: `https://api.decosa.ai`
- Health check: `GET https://api.decosa.ai/healthz`.
- **Hosted use is for made-up reviews only.** Real reviews from patients belong on a self-hosted box (see the self-host prompt). Say so wherever this is wired in.
- It never posts a reply and never stores the review. The owner reads every reply before posting it; the checks are English only.
## Auth: API key (or a demo session)
1. Preferred: an API key (`dk_…`) from "Get an API key" on the tool page. Keep it in `DECOSA_API_KEY`, never in code. Send `Authorization: Bearer $DECOSA_API_KEY`.
2. Without a key: `POST https://api.decosa.ai/demo/session` with `{"vertical": "review-reply"}` returns `{"token", "expires_at", "budget"}`. Demo sessions are limited per IP per hour; over a limit you get HTTP 429 with `Retry-After`. A token runs one request at a time (409 otherwise).
## Endpoints
- `POST /reviews/reply` (token). Body: `{review, rating (1-5), business_name, business_type?, healthcare?, reviewer_name?, contact?, facts?}`, or `{"sample_id": "..."}` for a sample from `GET /reviews/samples`. Add `"stream": true` (or `Accept: text/event-stream`) for events (`ready`, `receipt`, `draft`, `result`); the result carries `reply`, `source` (`model` or `fallback`), `drafts` with the blocked phrases, `record` and `usage` (tokens and cost at list price).
- `POST /reviews/check` (token): the same fields plus `reply`: the checks on a reply the owner wrote, with no model call.
- `GET /reviews/info`: what is checked, the HHS cases, limits, what is not checked.
- `POST /record/verify` (no auth): re-check the signed record.
Show the reply as a draft the owner must read, with the blocked phrases next to it. Never post it anywhere.
Run it yourself (containers)
On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.
# Decosa Reply to a review without breaking patient privacy: run it yourself (containers)
You are setting up Decosa's review-reply tool on this machine, so reviews never leave it. It returns:
- a short draft reply to a public review, ready for the owner to read and edit;
- the checks it passed, in code, and for a health or care business the patient-privacy rules plus a second check by the local model;
- the words it blocked, and a signed record that keeps only the review's fingerprint (SHA-256).
Nothing is sent to Decosa's hosted API.
Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with "not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.
Ask me before any command that needs sudo, and show me the command first.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/review-reply.zip (1 KB, 12 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py review-reply` (the api image carries the same bundle under /app/rehearsal/review-reply/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py review-reply --bundle review-reply.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the practice is treated as healthcare", "the final reply passes every check", "the reply never names the dentist"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin (docs.docker.com/engine/install), and the NVIDIA container toolkit; check `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file: `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`. Keep the `llm` service (Qwen3.8-27B on vLLM) and the `api` service; on `api` set `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b`; keep its data on a named volume; bind every port to 127.0.0.1.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check (the first start downloads the model weights).
4. Check: `curl -fsS http://127.0.0.1:<PORT>/reviews/info` answers; `GET /attest/signing-key` shows this box's public key.
5. Smoke test: get a token with `POST /demo/session {"vertical":"review-reply"}`, then `POST /reviews/reply {"sample_id": "dentist-named"}`. Expect:
- `.healthcare` is `true` and `.check.ok` is `true`;
- the reply contains none of `Alvarez`, `Priya`, `Morgan` or `crown`;
- every receipt has `"status": "attested"` and the record verifies at `POST /record/verify`;
- `POST /reviews/check` with the same review and the reply "See you at your next cleaning!" returns `"ok": false` with a `(patient privacy)` problem.
Finish with a summary and these reminders:
- Reviews from patients can hold health details: keep them on this machine; the model route stays local (`direct`).
- It never posts a reply and never stores the review. Read every reply before posting it; the checks are English only.
Run it on your own hardwareWhat it needs, and the prompt that sets it up
Run it on your own GPU
Same app, same pinned models, your hardware. Nothing goes to our servers and there are no Decosa charges.
Hardware check
Check your own hardware- CPU only, 64 GB RAMDoesn't fit
Qwen3.8-27B (NVIDIA NVFP4) needs a GPU.
- GeForce RTX 4090standard tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with A community 4-bit build of Qwen3.8-27B (AWQ or GGUF). This build is NVIDIA NVFP4, which needs a Blackwell GPU. (Memory is an estimate.)
- GeForce RTX 5090standard tierRuns
The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
- 2x GeForce RTX 5090standard tierRuns
The standard tier fits with changes: Split the language model across the GPUs with tensor parallelism (vLLM --tensor-parallel-size).
- L40Sstandard tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.
- H100 80 GB (SXM)standard tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B official FP8. This build is NVIDIA NVFP4, which needs a Blackwell GPU.
- RTX PRO 6000 Blackwell 96 GBstandard tierRuns
The standard tier fits (57.6 of 96 GB).
- 2x RTX PRO 6000 Blackwell 96 GBbest tierRuns
The standard tier fits (57.6 of 192 GB). The best tier fits too.
- Apple M3 Ultra (Mac Studio), 96 GBstandard tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon.
- Apple M5 Max, 64 GBstandard tierRuns
The standard tier fits with changes: Replace Qwen3.8-27B (NVIDIA NVFP4) with Qwen3.8-27B MLX 4-bit. MLX build for Apple Silicon.
Memory per component comes from measured footprints, the tool's stack.json, or an estimate from its parameter count, and each is labelled that way below. Only an RTX PRO 6000 and an M3 Ultra Mac Studio have actually been run.
On request. The container images and the compose file aren’t public yet. Ask for self-host access and Decosa sends the registry (DECOSA_REGISTRY) and the compose file’s URL (DECOSA_COMPOSE_URL) these steps use. They are the steps we tested end to end on a fresh machine.
- 1
Check the GPU, Docker and the NVIDIA Container Toolkit
The driver must see the GPU, and Docker must be able to pass it into a container.
nvidia-smi docker compose version docker run --rm --gpus all ubuntu nvidia-smi
- 2
Fetch the compose file
One file describes the API and the language model as services.
mkdir -p ~/decosa && cd ~/decosa curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml - 3
Pull and start
The first start downloads pinned model weights, tens of gigabytes.
docker compose pull docker compose up -d
- 4
Check health
Wait until the API reports ok with the language model loaded. Then point your app at the local base URL.
curl -fsS http://localhost:<PORT>/healthz # {"ok": true, "llm": true, ...} curl -fsS -X POST http://localhost:<PORT>/demo/session \ -H 'Content-Type: application/json' -d '{"vertical":"review-reply"}'
Set up with a coding agent, rehearse on mock data, then go private
- Set up with a coding agent. Paste the self-host prompt into a coding agent on the machine that will run the service. We recommend Claude Code with Claude Opus 5.5; any capable coding agent works.
- Rehearse on mock data. The agent runs the tool on a bundle of synthetic inputs and checks each answer against the bundle's
expected.json. Every check must print PASS. - Go private. Only then do you run your own data against the local API, yourself, on that machine. Never give the agent real data during setup: a coding agent that runs in the cloud sees everything in its context, so keep real data out of the chat and out of the files it reads.
docker compose exec api python scripts/rehearse.py review-reply
Download the mock-data bundle (1 KB, 12 checks)expected.json
A synthetic review for a made-up dental practice. The reply must pass every check, never confirm the reviewer is a patient, never name the dentist, the staff member or the reviewer, never mention the crown or the insurance; every model call has a signed receipt; the signed record verifies and a tampered one fails.
What the rehearsal checks
- the practice is treated as healthcare
- the final reply passes every check
- the reply never names the dentist
- the reply never names the staff member
- the reply never uses the reviewer's name
- the reply never mentions the crown
- the reply is signed off with the practice name
- an owner's reply that confirms the visit is refused
- the refusal names patient privacy
- the signed record verifies
- a record with its source changed no longer verifies
- every model call has a signed receipt
Licence: Synthetic: the practice, people and review are invented. Part of decosa-api.
Prompt for your coding agent
# Decosa Reply to a review without breaking patient privacy: run it yourself (containers)
You are setting up Decosa's review-reply tool on this machine, so reviews never leave it. It returns:
- a short draft reply to a public review, ready for the owner to read and edit;
- the checks it passed, in code, and for a health or care business the patient-privacy rules plus a second check by the local model;
- the words it blocked, and a signed record that keeps only the review's fingerprint (SHA-256).
Nothing is sent to Decosa's hosted API.
Status: the container images (${DECOSA_REGISTRY}/decosa-*) and the compose file are on request while self-host is in early access (not on a public registry yet): ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry as DECOSA_REGISTRY, the compose file URL as DECOSA_COMPOSE_URL, and pull access. If a pull fails with "not found", "denied" or "unauthorized", stop and tell me. Do not substitute other images.
Ask me before any command that needs sudo, and show me the command first.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/review-reply.zip (1 KB, 12 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py review-reply` (the api image carries the same bundle under /app/rehearsal/review-reply/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py review-reply --bundle review-reply.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the practice is treated as healthcare", "the final reply passes every check", "the reply never names the dentist"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## Steps
1. Docker: if `docker compose version` fails, install Docker Engine and the compose plugin (docs.docker.com/engine/install), and the NVIDIA container toolkit; check `docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi`.
2. Fetch the compose file: `mkdir -p ~/decosa && cd ~/decosa && curl -fsSL "${DECOSA_COMPOSE_URL}" -o compose.yaml`. Keep the `llm` service (Qwen3.8-27B on vLLM) and the `api` service; on `api` set `DECOSA_LLM_ROUTE=direct`, `DECOSA_LLM_URL=http://llm:8000/v1`, `DECOSA_LLM_MODEL=qwen3.8-27b`; keep its data on a named volume; bind every port to 127.0.0.1.
3. Pull and start: `docker compose pull && docker compose up -d`. Wait for the `llm` health check (the first start downloads the model weights).
4. Check: `curl -fsS http://127.0.0.1:<PORT>/reviews/info` answers; `GET /attest/signing-key` shows this box's public key.
5. Smoke test: get a token with `POST /demo/session {"vertical":"review-reply"}`, then `POST /reviews/reply {"sample_id": "dentist-named"}`. Expect:
- `.healthcare` is `true` and `.check.ok` is `true`;
- the reply contains none of `Alvarez`, `Priya`, `Morgan` or `crown`;
- every receipt has `"status": "attested"` and the record verifies at `POST /record/verify`;
- `POST /reviews/check` with the same review and the reply "See you at your next cleaning!" returns `"ok": false` with a `(patient privacy)` problem.
Finish with a summary and these reminders:
- Reviews from patients can hold health details: keep them on this machine; the model route stays local (`direct`).
- It never posts a reply and never stores the review. Read every reply before posting it; the checks are English only.
Help me customise for my hardware
Pick your GPU or Mac, or enter its memory. You get the tier that fits, the model swaps it needs, measured speed where we have it, and a setup prompt with those choices written in.
GeForce RTX 5090: 32 GB GDDR7, 1,792 GB/s, FP8 and NVFP4. NVIDIA product page
RunsReview reply with patient privacy on GeForce RTX 5090: use the Standard · the hosted demo, one 96 GB card tier
The standard tier fits with changes: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
Standard · the hosted demo, one 96 GB card: what changesuses estimates
- Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions.
Memory per component
- Drafts the reply: Qwen3.8-27B (NVIDIA NVFP4). ~57.6 GB (at least ~28 GB), weights 21.4 GB (from stack.json). Qwen3.8-27B NVFP4: Weights 19.9 GiB (21.4 GB), measured (field stack.json). The compose file gives the server 0.60 of a 96 GB card (57.6 GB) so the rest is FP8 KV cache for several sessions. The 28 GB minimum is an estimate: weights plus a short-context KV cache, which is why several stacks list a 32 GB RTX 5090 as 'estimate'. (stack.json lists 57 GB for this component.)
Expected speed
Not measured.
Not measured on this hardware. The only measured setups are an RTX PRO 6000 Blackwell and a Mac Studio M3 Ultra.
Setup prompt for this hardware
The self-host prompt for Review reply with patient privacy, with a hardware plan for GeForce RTX 5090 added after Step 0. Loading the full prompt; until then it points your agent at the prompt's URL.
# Set up Review reply with patient privacy on my hardware Fetch https://decosa.ai/prompts/review-reply-selfhost.md and follow it (including Step 0: rehearse on mock data first), with the hardware plan below applied. ## Hardware plan for this machine (from https://decosa.ai/self-host/hardware?use=review-reply) Target machine: GeForce RTX 5090 (32 GB of GPU memory; CUDA, FP8 and NVFP4). Quality tier: Standard · the hosted demo, one 96 GB card (standard). Fit check: runs with changes, about 28 GB of 32 GB used; some memory numbers are estimates, not measurements. First, check the machine: run `nvidia-smi` (or `rocm-smi`, or `sysctl hw.memsize` on a Mac) and confirm the GPUs and free memory match the line above. If they do not, stop and tell me before pulling anything. Use these components (the setup below describes the standard tier; change it to match): - Drafts the reply: Qwen3.8-27B (NVIDIA NVFP4) (nvidia/Qwen3.8-27B-NVFP4), 57.6 GB. Change: Qwen3.8-27B (NVIDIA NVFP4): run it at its smallest setting (about 28 GB instead of 57.6 GB), with a shorter context and fewer parallel sessions. GPU placement (set each service's device and its vLLM --gpu-memory-utilization to about the share shown): - GPU 0: Qwen3.8-27B (NVIDIA NVFP4) ~28 GB (88%); about 4 GB left During the rehearsal, watch GPU memory. If a model fails to load or runs out of memory, lower its --max-model-len and --max-num-seqs first, then its memory share, and tell me what you changed. The stack's own component list and compose layout: https://decosa.ai/prompts/review-reply-assemble.md
The proof
How we tested itEval results and end-to-end checks, hosted and self-hosted, with dates
Verified end to end
Hosted: verified 29 Sep 2026 · measured 29 Sep 2026: · p50 1.8 s · p95 4.2 s (100 runs) · ~<$0.001 per run · 1.55 receipts
Loading the nightly status…
Self-host: not yet verified
Measured cost to run: about $0.044 per 100 replies (hosted, 29 Sep 2026). Self-hosting is free: the code is open and the models are open-weight. You pay only for your own hardware and power.
Known limits (4)
- Timings and cost were measured on the pre-release server through the production gateway. On production (30 Sep 2026) the samples, a made-up review typed in by hand and the own-reply check were run end to end in a browser; self-hosting from the assemble prompt has not been verified yet.
- Reviews are synthetic, written by the same model family that drafts the replies; real reviews are messier.
- English only.
- A new phrasing that confirms a patient can still slip past both checks; read every reply before posting.
How it's builtThe steps, the models and what each one checks
Get an API key
- Call the review reply with patient privacy API from your own code in minutes.
- Every model answer carries a signed receipt.
- Nothing to install; we run the models.
Run it yourself, on request
- The same open models and app, on 1x RTX PRO 6000 (96 GB) for Qwen3.8-27B; the checks run on CPU.
- Data never leaves your machines, and there are no Decosa charges.
- One prompt for Claude Code or Codex assembles the whole stack.
- Early access: the container images are not public yet and the source needs access; the prompt says how to ask.
A short, warm reply to a public review, checked in code before you see it: no offers, no invented facts, and for a health or care practice never a word that confirms the reviewer is a patient.
For dentists, clinics, therapists and the office managers and agencies who answer their reviews, and for any local business that wants a reply that doesn't argue or promise things. Paste the review. The open model drafts a reply; code checks it for offers, contact details or numbers you didn't give, placeholders and arguing, and for a health or care business for anything that confirms a patient relationship: a visit, a treatment, a wait, a bill, a family member, the reviewer's or a provider's name. For those businesses a second, receipted model check reads the reply on its own. A draft that fails is redrafted once with the blocked words named; if it fails again you get a fixed safe reply. HHS's Office for Civil Rights has settled with dental practices and a psychiatric practice over review replies that disclosed patient information. Nothing is posted anywhere and nothing is stored.
- Deployment
- Hosted or self-host
- Regulatory
- HIPAA Privacy Rule, 45 CFR 164.502(a): a covered entity may not use or disclose protected health information except as permitted; confirming in a public reply that someone is a patient is a disclosure. HHS OCR resolutions over online review replies (read 29 Sep 2026 on hhs.gov): Elite Dental Associates, 2019, $10,000; New Vision Dental, $23,000; Manasa Health Center, $30,000, replies to negative Google reviews (hhs.gov/hipaa/for-professionals/compliance-enforcement/agreements: elite, new-vision, manasa). The FTC's rule on consumer reviews and testimonials (16 CFR Part 465, in force 21 Oct 2024) bars buying or suppressing reviews; this tool only drafts replies and never asks for a rating change. It checks common failure patterns; it is not legal advice.
Text description
A review, the stars, the business name and type go in. The open model drafts a reply; code checks it (offers, contact details, numbers, placeholders, arguing, patient privacy); for a health or care business a second model call reads the reply alone. A blocked draft is redrafted once, then a fixed safe reply. Out come the reply, what was blocked and a signed record that keeps only the review's fingerprint.
At a glance
- What it checks
- Offers and refunds, contact details or links you didn't give, numbers not in the review or your facts, placeholders, arguing with the reviewer, and that negative reviews are taken offline. For health and care businesses: anything that confirms a patient, a visit, a treatment, a condition, a wait, a bill or a family member's care, and any provider's or the reviewer's name.
- What it never does
- Post a reply, ask for a rating change, or store the review.
- Data retention
- None: the review and the reply live in memory for the request. The signed record holds the review's SHA-256, never its text.
- What leaves the box (hosted demo)
- The review goes to the open model on Decosa's hosted service through the receipted gateway. Nothing goes to a third party.
- Model calls per reply
- 1.55 on average on the held-out set (a draft, sometimes a redraft, and for health and care businesses a privacy check of each draft that passes the code)
- Cost per reply
- A fraction of a cent per reply on average at list price, held-out set.
- Also used in
- ElmoSEO's review replies run the same guard (the TypeScript kit).
Pick the tier for the quality you need
Same app at every tier. What changes is the models, the hardware they need, and whether receipts are signed. Scores are measured with the source named, or marked not measured.
Lite
one 48 GB card
The same guard and flow on a smaller mixture-of-experts model. Faster and cheaper; accuracy on this task unknown.
- Models
- @decosa/site-kit review guard
- Gemma 4 26B A4B (instruction-tuned)
- Hardware
- 1x L40S or RTX 6000 Ada 48 GB (not measured)
- Quality evidence
- accuracy on this tasknot measured yet
- Latency
- not measured yet
- Verification
- Proof: partialSelf-host onlyDirect route: calls are attested by the box's key; no gateway receipts.
- In the hosted demo
Standard
the hosted demo, one 96 GB card
Qwen3.8-27B drafts; the guard is code; health and care replies also get the model's privacy check. One to four model calls per reply, each receipted.
- Models
- @decosa/site-kit review guard
- Qwen3.8-27B (NVIDIA NVFP4)
- Hardware
- 1x RTX PRO 6000 Blackwell 96 GB
- Quality evidence
- held-out healthcare replies that confirm a patient (blind judge)0 of 40decosa-api docs/evals/review-reply.md, gateway route, 29 Sep 2026
- held-out replies a business could post as written (blind judge)93 of 100decosa-api docs/evals/review-reply.md
- Latency
- measured on the held-out set, shared gateway: seconds per reply
- Verification
- Proof: strongHosted: gateway-signed receipt per model call. Self-hosted: attested by the box's key.
Best
DeepSeek-V4-Flash on two more cards
A larger model for long, multi-claimant files.
- Models
- @decosa/site-kit review guard
- DeepSeek-V4-Flash (NVIDIA NVFP4)
- Hardware
- 2x RTX PRO 6000 96 GB
- Quality evidence
- accuracy on this tasknot measured yet
- Latency
- not measured yet
- Verification
- Proof: partialSelf-host onlyNot a hosted model: calls are attested by the box's key only.
- Needs more compute
Wanted: the best setup
two large judges from different families
DeepSeek-V4-Flash and GLM-5.3-Flash each read the file, and a finding stands when they agree; disagreements go to the reviewer. Claim files stay on your own hardware, never on community providers. Not served yet.
- Models
- DeepSeek-V4-Flash (NVIDIA NVFP4)
- GLM-5.3-Flash
- Hardware
- Your own hardware: 2x 96 GB cards for DeepSeek-V4-Flash plus 2x 96 GB for GLM-5.3-Flash, or one Mac Studio with 512 GB holding both 4-bit builds (156 + 165 GB, sizes from our Mac; not run together yet). Estimate.
- Quality evidence
- accuracy on this tasknot measured yet
- Latency
- not measured yet
- Verification
- No proof yetSelf-host onlyOn your own hardware its calls are attested by the box's key only: not a hosted model there, so no gateway receipts. Never sent to community providers.
Not served yet. It needs more than one 96 GB card, so it runs on your own bigger box.
Every model in the stack
| Model | Tiers | Params · VRAM | Verification | Details |
|---|---|---|---|---|
| After the session | ||||
The review-reply guard: offers, contact details and numbers not given, placeholders, arguing, and the patient-privacy rules. Deterministic code, the same in TypeScript and Python.@decosa/site-kit review guard 0 GBNo proof yetIn the hosted demo | LiteStandardBest | 0 GB | No proof yetIn the hosted demo | |
| ||||
Drafts the reply (and a redraft when the code checks block the first); for health and care businesses, a second call reads the reply alone and says whether it confirms a patient. The checks themselves are code (the open-source site kit).Qwen3.8-27B (NVIDIA NVFP4)nvidia/Qwen3.8-27B-NVFP4 on Hugging Face (opens in a new tab) 27.8B · 57 GBProof: strongIn the hosted demo | Standard | 27.8B · 57 GB | Proof: strongIn the hosted demo | |
| ||||
Would draft the reply and run the privacy check; not measured on this task.Gemma 4 26B A4B (instruction-tuned)google/gemma-4-26B-A4B-it on Hugging Face (opens in a new tab) 25.2B (3.8B active)No proof yetSelf-host only | Lite | 25.2B (3.8B active) | No proof yetSelf-host only | |
| ||||
Would draft the reply and run the privacy check; not measured on this task.DeepSeek-V4-Flash (NVIDIA NVFP4)nvidia/DeepSeek-V4-Flash-NVFP4 on Hugging Face (opens in a new tab) 284B (13B active) · 192 GBNo proof yetSelf-host only | BestWanted | 284B (13B active) · 192 GB | No proof yetSelf-host only | |
| ||||
| Other | ||||
Second judge, from another familyGLM-5.3-Flashzai-org/GLM-5.3-Flash on Hugging Face (opens in a new tab) 321B (18B active) · about 170 GB (estimate)No proof yetSelf-host only | Wanted | 321B (18B active) · about 170 GB (estimate) | No proof yetSelf-host only | |
| ||||
Tools, services and hardware
Tools
The case the privacy rules are built around: replies that disclosed patients' names and care.
- @decosa/site-kit (opens in a new tab)Apache-2.0
The guard, in TypeScript; the Python mirror runs in decosa-api. Shared fixtures keep them identical.
Services
- decosa-api:8445
${DECOSA_REGISTRY}/decosa-api:0.1.0The guard, the drafting prompt, the redraft and fallback, signing and the HTTP API (/reviews/*). No GPU. Binds 127.0.0.1 by default.
- decosa-llm:8000
${DECOSA_REGISTRY}/decosa-llm:0.1.0vLLM OpenAI endpoint for Qwen3.8-27B. Internal to the compose network.
Hardware
- 1x RTX PRO 6000 Blackwell 96 GB Fits
Measured: the hosted demo's Qwen3.8-27B runs on one of these cards on our server.
- 1x L40S / RTX 6000 Ada 48 GB
Not measured. FP8 Qwen3.8-27B with a shorter context, or Gemma 4 26B A4B (lite).
- CPU only Fits
The date rules, the sums, the header checks, signing and verification need no GPU; reading the text needs the model.
Latency per lane
- one reply, health or care business, busy shared gateway2.2 s
Measureddecosa-api docs/evals/review-reply.md, held-out set, gateway route, 29 Sep 2026 (p95 4.6 s)
- one reply, any business, busy shared gateway1.8 s
Measureddecosa-api docs/evals/review-reply.md, held-out set, gateway route, 29 Sep 2026 (p95 4.237 s)
- check your own reply (no model)50 ms
MeasuredPOST /reviews/check: code only
Notes
- English only: a review in another language gets an English reply, and the guard's word lists are English.
- The guard blocks broadly on purpose: many healthcare drafts end in the fixed safe reply, which is generic but safe.
Run this exact stack on your machine
Paste into Claude Code / Codex to assemble this stack locally. The prompt checks your GPU, pulls the pinned models, writes the compose file and runs a smoke test.
# Assemble Decosa's "Reply to a review without breaking patient privacy" on this machine
You are setting up the Decosa review-reply tool on this Linux machine for a dental, medical or therapy practice (or an agency that answers reviews for practices). It returns:
- a short draft reply to a public review, in the owner's voice, ready to read and edit;
- the checks it passed, in code: no offers, no contact details or numbers the owner didn't give, no placeholders, not defensive, negative reviews taken offline;
- for a health or care business, the patient-privacy rules: nothing that confirms the reviewer is a patient or mentions a visit, treatment, condition, wait, bill, family member, provider or the reviewer's name, plus a second check of the reply by the local model;
- the exact words it blocked, and a signed record that keeps only the review's fingerprint (SHA-256).
Work step by step, show me each command before running anything that needs sudo, and stop if a check fails.
**Before anything else, remind me:**
- Reviews from patients can hold health details: keep them on this machine; the model route stays local (`direct`).
- It never posts a reply and never stores the review. Read every reply before posting it; the checks catch common patterns, not every possible phrasing, and they are English only.
Repeat both points in your final summary.
## Step 0: set up with a coding agent, rehearse on mock data, then go private
This prompt is for a coding agent running on the machine that will host the service. We recommend Claude Code with
Claude Opus 5.5; any capable coding agent works. Work in this order:
1. Set up on mock data only. During the whole setup you (the agent) work with the synthetic sample bundle below and
nothing else. Do not ask me for real data, and do not open, read, list or copy files that hold real data, even to
"test with something realistic".
2. Rehearse. When the steps below are done and the service is healthy, fetch the mock-data bundle for this tool,
https://decosa.ai/samples/review-reply.zip (1 KB, 12 checks, synthetic or openly licensed: see `licence` in expected.json),
show me what is in it, and run the rehearsal against the local API:
`docker compose exec api python scripts/rehearse.py review-reply` (the api image carries the same bundle under /app/rehearsal/review-reply/;
with no key set, the script asks the local API for a short demo token). From a decosa-api checkout instead:
`python scripts/rehearse.py review-reply --bundle review-reply.zip --base-url http://127.0.0.1:<PORT>`.
It sends the mock inputs to the local API and prints PASS or FAIL for each expected property (for example: "the practice is treated as healthcare", "the final reply passes every check", "the reply never names the dentist"). Show me
the full output. Every check must pass. If one fails, fix the install and run it again; never edit `expected.json`
to make a check pass.
3. Stop there. Once the rehearsal passes, tell me, and I will run my own data against the local API myself, on this
machine.
For the person running this: a coding agent that runs in the cloud sees everything in its context, including files it
reads, command output and anything pasted into the chat. Keep real data out of the chat and out of anything the agent
can read. Switch to your own data only after the rehearsal has passed and the agent's work is done.
## What you are building
| service | image | model | port |
|---|---|---|---|
| `llm` | `${DECOSA_REGISTRY}/decosa-llm:0.1.0` (vLLM 0.29.0, `vllm/vllm-openai@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1`) | `nvidia/Qwen3.8-27B-NVFP4` @ `482ca0f3832238542f8f5295dde86b5f22711d80`, Apache-2.0 | internal 8000 |
| `api` | `${DECOSA_REGISTRY}/decosa-api:0.1.0` (no GPU) | none | `127.0.0.1:8445` |
## 1. Check the GPU, driver and Docker
1. Run `nvidia-smi`. I need one NVIDIA GPU with at least 48 GB and driver 580 or newer.
- Blackwell (RTX PRO 6000, B200): use the defaults below (NVFP4). This is the measured setup.
- Hopper (H100/H200) or 48 GB Ada/L40S: set `LLM_MODEL=Qwen/Qwen3.8-27B-FP8` and `LLM_REVISION=main`. On a 48 GB card, also set `LLM_MAX_LEN=32768`. Not measured.
- Under 48 GB: stop and tell me it will not fit. (`POST /reviews/check`, which checks a reply you wrote, needs no GPU: the `api` service alone runs it.)
2. Check `docker --version`, `docker compose version` and `docker run --rm --gpus all ubuntu nvidia-smi`. If Docker or the NVIDIA Container Toolkit is missing, install them from the official Docker and NVIDIA repositories, then run `sudo nvidia-ctk runtime configure --runtime=docker` and restart Docker.
3. Confirm about 60 GB of free disk.
## 2. Get the images
The images are **on request** while self-host is in early access: ask at https://decosa.ai/contact?topic=self-host, and Decosa sends the registry (set it as `DECOSA_REGISTRY`), pull access and the compose file. Try `docker pull ${DECOSA_REGISTRY}/decosa-{llm,api}:0.1.0`.
If a pull fails, build from source once the `decosa-api` source is published: in the clone, `docker build -f docker/api/Dockerfile -t ${DECOSA_REGISTRY}/decosa-api:0.1.0 .`, and `docker compose build llm` from its compose file. If neither works, stop and tell me.
## 3. Write the compose file
Create `~/decosa-reviewreply/.env`:
```bash
DECOSA_TAG=0.1.0
DECOSA_GPU=0
LLM_MODEL=nvidia/Qwen3.8-27B-NVFP4
LLM_REVISION=482ca0f3832238542f8f5295dde86b5f22711d80
LLM_MAX_LEN=65536
LLM_GPU_UTIL=0.85
DECOSA_SIGNER_NAME="<who signs these records, e.g. Example Agency>"
```
Create `~/decosa-reviewreply/docker-compose.yml` with exactly these services:
```yaml
name: decosa-reviewreply
x-health: &health
interval: 15s
timeout: 5s
retries: 5
services:
llm:
image: ${DECOSA_REGISTRY}/decosa-llm:${DECOSA_TAG}
deploy: { resources: { reservations: { devices: [ { driver: nvidia, device_ids: ["${DECOSA_GPU:-0}"], capabilities: [gpu] } ] } } }
ipc: host
restart: unless-stopped
volumes: [hf-cache:/root/.cache/huggingface]
command: ["${LLM_MODEL}", "--revision", "${LLM_REVISION}", "--served-model-name", "qwen3.8-27b",
"--language-model-only", "--max-model-len", "${LLM_MAX_LEN}", "--gpu-memory-utilization", "${LLM_GPU_UTIL}",
"--max-num-seqs", "16", "--kv-cache-dtype", "fp8_e4m3", "--speculative-config", '{"method":"mtp","num_speculative_tokens":3}',
"--seed", "0", "--enable-force-include-usage", "--disable-uvicorn-access-log", "--host", "0.0.0.0", "--port", "8000"]
healthcheck: { <<: *health, test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8000/health', timeout=4)"], start_period: 900s }
api:
image: ${DECOSA_REGISTRY}/decosa-api:${DECOSA_TAG}
restart: unless-stopped
depends_on: { llm: { condition: service_healthy } }
environment:
DECOSA_LLM_ROUTE: direct # local model only; receipts are signed by this box's key ("attested")
DECOSA_LLM_URL: http://llm:8000/v1
DECOSA_LLM_MODEL: qwen3.8-27b
DECOSA_LOCAL_SIGNING: "on" # Ed25519 key created at /data/attest/ed25519.pem on first start
DECOSA_SIGNER_NAME: ${DECOSA_SIGNER_NAME}
DECOSA_SESSIONS_PER_IP_HOUR: "1000"
DECOSA_BUDGET_LLM_TOKENS: "200000"
DECOSA_SESSION_TTL_S: "28800"
DECOSA_CORS_ORIGIN_REGEX: '^https?://(localhost|127\.0\.0\.1)(:\d+)?$$'
DECOSA_REVIEWREPLY_MAX_CONCURRENT: "6"
ports: ["127.0.0.1:8445:8445"]
volumes: [decosa-data:/data]
healthcheck: { <<: *health, test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8445/healthz', timeout=4)"], start_period: 20s }
volumes: { hf-cache: {}, decosa-data: {} }
```
Use the named volume `decosa-data` exactly as written; a host bind mount owned by root makes the API fail on `/data/keys.sqlite`.
Run `docker compose up -d`, then poll `docker compose ps` until both services are healthy (5-10 minutes the first time). `curl -s localhost:8445/healthz` should show `"llm": true`.
## 4. Smoke test
```bash
API=localhost:8445
TOKEN=$(curl -s $API/demo/session -H 'content-type: application/json' -d '{"vertical":"review-reply"}' | jq -r .token)
curl -s $API/reviews/reply -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' \
-d '{"sample_id":"dentist-named"}' > /tmp/out.json
jq '{reply, source, attempts, healthcare, ok: .check.ok}' /tmp/out.json
jq '{record}' /tmp/out.json | curl -s $API/record/verify -H 'content-type: application/json' -d @- | jq '{ok, summary}'
curl -s $API/reviews/check -H "authorization: Bearer $TOKEN" -H 'content-type: application/json' \
-d '{"review":"Dr. Alvarez fixed my crown.","rating":5,"business_name":"Juniper Street Dental","business_type":"Dentist","reply":"So glad Dr. Alvarez fixed your crown! See you at your next cleaning."}' | jq '{ok, problems}'
```
The sample `dentist-named` is synthetic. Pass if:
- `.healthcare` is `true` and `.check.ok` is `true`;
- the reply contains none of `Alvarez`, `Priya`, `Morgan` or `crown`, and ends with `– Juniper Street Dental`;
- every receipt has `"status": "attested"` and the record verifies;
- the last command returns `"ok": false` with problems ending in `(patient privacy)`.
## 5. Point the app at the local API
- Base URL: `http://localhost:8445`. For a web app, set `NEXT_PUBLIC_DECOSA_API=http://localhost:8445`; add other origins to `DECOSA_CORS_ORIGINS`.
- `POST /reviews/reply` takes `{review, rating (1-5), business_name, business_type?, healthcare?, reviewer_name?, contact?, facts?}`. It returns JSON, or streams Server-Sent Events with `Accept: text/event-stream`. `healthcare` overrides the detection from `business_type`.
- `POST /reviews/check` takes the same fields plus `reply` and runs the checks on a reply you wrote, with no model call.
- `GET /reviews/info` lists what is checked, the HHS cases, the limits and what is not checked.
- The server stores nothing. Anyone can re-check a signed record with `POST /record/verify` against the key at `GET /attest/signing-key`. No call leaves the box.
- The checks are also a TypeScript library, `@decosa/site-kit` (Apache-2.0, in `clients/site-kit` of the decosa-api source), for checking replies inside your own app.
- Keep the API on 127.0.0.1. For other users on the LAN, put a TLS reverse proxy with authentication in front and set `DECOSA_TRUSTED_PROXIES`.
## 6. Hosted routes (off, and leave them off)
`DECOSA_LLM_ROUTE=gateway` would send prompts, which contain the reviews, to a hosted gateway. Never use it for real patients' reviews.
Finish with a summary: what is running, the health output, the smoke-test results, and the reminders above.Technical detailsModels, where it runs, labels
In short
Last reviewed
- What it is
- A hipaa review response has to thank or apologise without saying the reviewer was ever a patient. This drafts the reply and checks it in code before you see it.
- Who it's for
- Dentists, clinics, therapists and the office managers and agencies who answer their reviews.
- Where it runs
- Hosted demo on made-up reviews; self-host for your practice's real reviews
- Key numbers
On a held-out set judged blind, no healthcare reply confirmed a patient; the price is a generic fixed reply for many of them.
- 0 / 40 Healthcare replies that confirm a patient (blind judge) (test split, n = 40)
- 93 / 100 Replies a business could post as written (blind judge) (test split, n = 100)
- 24 / 40 Healthcare replies that fell back to the fixed safe reply (test split, n = 40)
- 1.8 s Median end-to-end run, hosted (QA sweep 2026-09-29)
- Models
- Qwen3.8-27B (drafts, and reads health and care replies for patient privacy); the checks are code from the open-source site kit
- Where
- Hosted demo on made-up reviews; self-host for your practice's real reviews
- Checks
- Checks in code (Apache-2.0 site kit); receipt per model call; signed record with the review's fingerprint only
- Industry
- Healthcare · Sales and marketing
- Input
- Text and documents
- Output
- Notes, reports and drafts · Signed record or verdict
- Data
- Personal data
- Hardware
- 1× 96 GB GPU
- Licence
- Permissive (Apache-2.0, MIT)
- Runs in
- Decosa hosted · Self-host
- Built from
- Site checks · Signed record
Questions people ask
How do I write a hipaa review response?
Thank or apologise in general terms, never confirm the reviewer is a patient or mention their visit, treatment, wait or bill, and invite them to contact the office. This tool drafts that and blocks the phrases that break it.
Does it post the reply?
No. It gives you a draft to read and edit. Nothing is posted anywhere, and the review is not stored.
What if the model writes something that confirms a patient?
Code checks every draft, and for health and care businesses a second model check reads it too. A blocked draft is redrafted once; if it fails again you get a fixed safe reply.
Can I check a reply I wrote myself?
Yes: POST /reviews/check runs the same checks on your own reply with no model call, and the guard is open source (Apache-2.0) in the site kit.
Ask a question or leave feedbackWe read every message and publish useful answers
Ask about Review reply with patient privacy
We read every message. Questions, comments and our answers show here once we have reviewed and approved them.
Loading questions…