{"schema_version":"1","site":"https://decosa.ai","id":"review-reply","num":"182","name":"Review reply with patient privacy","tool_name":"Reply to a review without breaking patient privacy","short":"Review reply","blurb":"Paste a Google or Yelp review. The open model drafts a short, warm reply, and code checks it before you see it: no offers or refunds in public, no contact details or numbers you didn't give, no arguing, and for a dentist, clinic or therapist never a word that confirms the reviewer is a patient: no visit, treatment, wait, bill, family member or name. A second model check reads health and care replies on their own. A blocked draft is redrafted once, then a fixed safe reply. Nothing is posted and the review isn't stored.","status":"live","labels":{"industry":["healthcare","sales-marketing"],"job":["draft","review"],"input":["text"],"deploy":["hosted","selfhost"],"status":"live","output":["text","record"],"data":["pii"],"hardware":"gpu-96","licence":"permissive"},"industries":["healthcare","sales-marketing"],"runs_in":["hosted","selfhost"],"part_of":[],"built_from":["site-checks","signed-record"],"models":"Qwen3.8-27B (drafts, and reads health and care replies for patient privacy); the checks are code from the open-source site kit","where":"Hosted demo on made-up reviews; self-host for your practice's real reviews","hardware":"1x RTX PRO 6000 (96 GB) for Qwen3.8-27B; the checks run on CPU","final_artifact":"A reply to read and post yourself, the words that were blocked, and a signed record that keeps only the review's fingerprint.","self_host_first":false,"verification":{"receipt_coverage":"full","summary":"Checks in code (Apache-2.0 site kit); receipt per model call; signed record with the review's fingerprint only","manual_qa":{"hosted":{"date":"2026-09-29","result":"pass","p50_ms":1801,"p95_ms":4237,"runs":100,"receipts_per_run":1.55,"cost_per_run_usd":0.000441},"selfhost":null,"known_limits":["Timings and cost were measured on the pre-release server through the production gateway. On production (30 Sep 2026) the samples, a made-up review typed in by hand and the own-reply check were run end to end in a browser; self-hosting from the assemble prompt has not been verified yet.","Reviews are synthetic, written by the same model family that drafts the replies; real reviews are messier.","English only.","A new phrasing that confirms a patient can still slip past both checks; read every reply before posting."],"nightly_covers":null},"nightly":"https://api.decosa.ai/verify/status"},"eval_summary":{"metrics":[{"name":"Healthcare replies that confirm a patient (blind judge)","value":"0 / 40","unit":null,"n":40,"split":"test","note":"held-out set generated after the guard was frozen; target 0"},{"name":"Replies a business could post as written (blind judge)","value":"93 / 100","unit":null,"n":100,"split":"test","note":"healthcare 38/40, other 55/60"},{"name":"Healthcare replies that fell back to the fixed safe reply","value":"24 / 40","unit":null,"n":40,"split":"test","note":"the price of blocking broadly"},{"name":"Replies with an offer or a fact nobody gave (blind judge)","value":"4 / 100","unit":null,"n":100,"split":"test","note":null},{"name":"First guard on the development set: healthcare replies that confirm a patient","value":"9 / 40","unit":null,"n":40,"split":"dev","note":"code rules only; led to the model privacy check and broader rules"}],"dataset":"200 synthetic reviews (80 for health and care businesses) written by Qwen3.8-27B from seeded plans: set 1 for development, set 2 held out and generated after the guard was frozen. Replies judged blind by Claude Code (Opus), which saw only the business type, stars, review and reply.","held_out":true,"caveats":["One judge (a frontier model), no human labels.","Synthetic reviews written by the same model family that drafts the replies.","English only.","The guard blocks broadly, so many health and care replies are the generic fixed reply."],"date":"2026-09-29","doc_url":"https://decosa.ai/metrics/evals/review-reply"},"stack":{"summary":"For dentists, clinics, therapists and the office managers and agencies who answer their reviews, and for any local business that wants a reply that doesn't argue or promise things. Paste the review. The open model drafts a reply; code checks it for offers, contact details or numbers you didn't give, placeholders and arguing, and for a health or care business for anything that confirms a patient relationship: a visit, a treatment, a wait, a bill, a family member, the reviewer's or a provider's name. For those businesses a second, receipted model check reads the reply on its own. A draft that fails is redrafted once with the blocked words named; if it fails again you get a fixed safe reply. HHS's Office for Civil Rights has settled with dental practices and a psychiatric practice over review replies that disclosed patient information. Nothing is posted anywhere and nothing is stored.","tagline":"A short, warm reply to a public review, checked in code before you see it: no offers, no invented facts, and for a health or care practice never a word that confirms the reviewer is a patient.","deployment":"hosted-or-self-host","regulatory_note":"HIPAA Privacy Rule, 45 CFR 164.502(a): a covered entity may not use or disclose protected health information except as permitted; confirming in a public reply that someone is a patient is a disclosure. HHS OCR resolutions over online review replies (read 29 Sep 2026 on hhs.gov): Elite Dental Associates, 2019, $10,000; New Vision Dental, $23,000; Manasa Health Center, $30,000, replies to negative Google reviews (hhs.gov/hipaa/for-professionals/compliance-enforcement/agreements: elite, new-vision, manasa). The FTC's rule on consumer reviews and testimonials (16 CFR Part 465, in force 21 Oct 2024) bars buying or suppressing reviews; this tool only drafts replies and never asks for a rating change. It checks common failure patterns; it is not legal advice.","components":[{"id":"guard","role":"The review-reply guard: offers, contact details and numbers not given, placeholders, arguing, and the patient-privacy rules. Deterministic code, the same in TypeScript and Python.","name":"@decosa/site-kit review guard","hf_repo":null,"license":"Apache-2.0","params":null,"quant":null,"vram_gb":0,"memory_gb_estimate":null,"engine":"Python (decosa-api) / TypeScript (@decosa/site-kit)","receipt_coverage":"none","in_hosted_demo":true,"tiers":["lite","standard","best"],"alternative_to":null},{"id":"llm","role":"Drafts the reply (and a redraft when the code checks block the first); for health and care businesses, a second call reads the reply alone and says whether it confirms a patient. The checks themselves are code (the open-source site kit).","name":"Qwen3.8-27B (NVIDIA NVFP4)","hf_repo":"nvidia/Qwen3.8-27B-NVFP4","license":"Apache-2.0","params":"27.8B","quant":"NVFP4 (MLP) + FP8 (attention/GDN), FP8 KV cache, MTP speculative decoding k=3","vram_gb":57,"memory_gb_estimate":null,"engine":"vLLM 0.29.0","receipt_coverage":"strong","in_hosted_demo":true,"tiers":["standard"],"alternative_to":null},{"id":"llm-lite","role":"Would draft the reply and run the privacy check; not measured on this task.","name":"Gemma 4 26B A4B (instruction-tuned)","hf_repo":"google/gemma-4-26B-A4B-it","license":"Apache-2.0 (model card also links the Gemma 4 licence page)","params":"25.2B","quant":"BF16 weights; FP8 at load time (vLLM --quantization fp8) to fit a 48 GB card","vram_gb":null,"memory_gb_estimate":null,"engine":"vLLM 0.29.0 (Gemma4ForConditionalGeneration is in its model registry)","receipt_coverage":"none","in_hosted_demo":false,"tiers":["lite"],"alternative_to":null},{"id":"llm-best","role":"Would draft the reply and run the privacy check; not measured on this task.","name":"DeepSeek-V4-Flash (NVIDIA NVFP4)","hf_repo":"nvidia/DeepSeek-V4-Flash-NVFP4","license":"MIT","params":"284B","quant":"NVFP4 experts + FP8 (about 159-176 GB of weights)","vram_gb":192,"memory_gb_estimate":null,"engine":"vLLM B12X community build, TP2 on 2x RTX PRO 6000, MTP draft fixed by the kit's patches","receipt_coverage":"none","in_hosted_demo":false,"tiers":["best"],"alternative_to":null},{"id":"glm-wanted","role":"Second judge, from another family","name":"GLM-5.3-Flash","hf_repo":"zai-org/GLM-5.3-Flash","license":"MIT","params":"321B","quant":"NVFP4 on NVIDIA (nvidia/GLM-5.3-Flash-NVFP4, about 170-186 GB, unconfirmed); MLX 4-bit on a Mac (165 GB)","vram_gb":null,"memory_gb_estimate":170,"engine":"SGLang SM120 build, TP2 on 2x 96 GB (vLLM is broken on sm_120 for this model, and the SGLang build hung on our server), or mlx-lm on a Mac with 192 GB or more","receipt_coverage":"none","in_hosted_demo":false,"tiers":["wanted"],"alternative_to":null}],"tiers":[{"id":"lite","label":"Lite · one 48 GB card","summary":"The same guard and flow on a smaller mixture-of-experts model. Faster and cheaper; accuracy on this task unknown.","components":["llm-lite"],"hardware":"1x L40S or RTX 6000 Ada 48 GB (not measured)","quality_evidence":[{"metric":"accuracy on this task","value":"not measured yet","source":null}],"latency_note":"not measured yet","in_hosted_demo":false,"receipt_coverage":"partial","receipt_note":"Direct route: calls are attested by the box's key; no gateway receipts.","hosting":null},{"id":"standard","label":"Standard · the hosted demo, one 96 GB card","summary":"Qwen3.8-27B drafts; the guard is code; health and care replies also get the model's privacy check. One to four model calls per reply, each receipted.","components":["llm"],"hardware":"1x RTX PRO 6000 Blackwell 96 GB","quality_evidence":[{"metric":"held-out healthcare replies that confirm a patient (blind judge)","value":"0 of 40","source":"decosa-api docs/evals/review-reply.md, gateway route, 29 Sep 2026"},{"metric":"held-out replies a business could post as written (blind judge)","value":"93 of 100","source":"decosa-api docs/evals/review-reply.md"}],"latency_note":"measured on the held-out set, shared gateway: seconds per reply","in_hosted_demo":true,"receipt_coverage":"strong","receipt_note":"Hosted: gateway-signed receipt per model call. Self-hosted: attested by the box's key.","hosting":null},{"id":"best","label":"Best · DeepSeek-V4-Flash on two more cards","summary":"A larger model for long, multi-claimant files.","components":["llm-best"],"hardware":"2x RTX PRO 6000 96 GB","quality_evidence":[{"metric":"accuracy on this task","value":"not measured yet","source":null}],"latency_note":"not measured yet","in_hosted_demo":false,"receipt_coverage":"partial","receipt_note":"Not a hosted model: calls are attested by the box's key only.","hosting":null},{"id":"wanted","label":"Wanted · two large judges from different families","summary":"DeepSeek-V4-Flash and GLM-5.3-Flash each read the file, and a finding stands when they agree; disagreements go to the reviewer. Claim files stay on your own hardware, never on community providers. Not served yet.","components":["llm-best","glm-wanted"],"hardware":"Your own hardware: 2x 96 GB cards for DeepSeek-V4-Flash plus 2x 96 GB for GLM-5.3-Flash, or one Mac Studio with 512 GB holding both 4-bit builds (156 + 165 GB, sizes from our Mac; not run together yet). Estimate.","quality_evidence":[{"metric":"accuracy on this task","value":"not measured yet","source":null}],"latency_note":"not measured yet","in_hosted_demo":false,"receipt_coverage":"none","receipt_note":"On your own hardware its calls are attested by the box's key only: not a hosted model there, so no gateway receipts. Never sent to community providers.","hosting":"own-hardware"}],"alternates":[],"services":[{"name":"decosa-api","port":8445,"image":"${DECOSA_REGISTRY}/decosa-api:0.1.0","purpose":"The guard, the drafting prompt, the redraft and fallback, signing and the HTTP API (/reviews/*). No GPU. Binds 127.0.0.1 by default."},{"name":"decosa-llm","port":8000,"image":"${DECOSA_REGISTRY}/decosa-llm:0.1.0","purpose":"vLLM OpenAI endpoint for Qwen3.8-27B. Internal to the compose network."}],"tools":[{"name":"HHS OCR: Elite Dental Associates resolution agreement","url":"https://www.hhs.gov/hipaa/for-professionals/compliance-enforcement/agreements/elite/index.html","license":"US government work","purpose":"The case the privacy rules are built around: replies that disclosed patients' names and care."},{"name":"@decosa/site-kit","url":"https://decosa.ai/contact?topic=self-host","license":"Apache-2.0","purpose":"The guard, in TypeScript; the Python mirror runs in decosa-api. Shared fixtures keep them identical."}],"hardware":[{"tier":"1x RTX PRO 6000 Blackwell 96 GB","fits":true,"notes":"Measured: the hosted demo's Qwen3.8-27B runs on one of these cards on our server."},{"tier":"1x L40S / RTX 6000 Ada 48 GB","fits":null,"notes":"Not measured. FP8 Qwen3.8-27B with a shorter context, or Gemma 4 26B A4B (lite)."},{"tier":"CPU only","fits":true,"notes":"The date rules, the sums, the header checks, signing and verification need no GPU; reading the text needs the model."}],"latency":[{"lane":"one reply, health or care business, busy shared gateway","typical_ms":2154,"source":"decosa-api docs/evals/review-reply.md, held-out set, gateway route, 29 Sep 2026 (p95 4.6 s)"},{"lane":"one reply, any business, busy shared gateway","typical_ms":1801,"source":"decosa-api docs/evals/review-reply.md, held-out set, gateway route, 29 Sep 2026 (p95 4.237 s)"},{"lane":"check your own reply (no model)","typical_ms":50,"source":"POST /reviews/check: code only"}],"benchmark":null,"notes":["English only: a review in another language gets an English reply, and the guard's word lists are English.","The guard blocks broadly on purpose: many healthcare drafts end in the fixed safe reply, which is generic but safe."]},"buyer_facts":[{"label":"What it checks","value":"Offers and refunds, contact details or links you didn't give, numbers not in the review or your facts, placeholders, arguing with the reviewer, and that negative reviews are taken offline. For health and care businesses: anything that confirms a patient, a visit, a treatment, a condition, a wait, a bill or a family member's care, and any provider's or the reviewer's name."},{"label":"What it never does","value":"Post a reply, ask for a rating change, or store the review."},{"label":"Data retention","value":"None: the review and the reply live in memory for the request. The signed record holds the review's SHA-256, never its text."},{"label":"What leaves the box (hosted demo)","value":"The review goes to the open model on Decosa's hosted service through the receipted gateway. Nothing goes to a third party."},{"label":"Model calls per reply","value":"1.55 on average on the held-out set (a draft, sometimes a redraft, and for health and care businesses a privacy check of each draft that passes the code)"},{"label":"Cost per reply","value":"A fraction of a cent per reply on average at list price, held-out set."},{"label":"Also used in","value":"ElmoSEO's review replies run the same guard (the TypeScript kit)."}],"data_handling":{"page":"/data#review-reply","self_host":{"level":"confidential","leaves":"nothing","summary":"Runs on your machine; nothing is sent to Decosa or a third party by default."},"hosted":{"level":"operator-processed","demo_only":false,"summary":"TLS to Decosa's server, then decrypted and processed by Decosa's API server, with the open models run by NEAR AI through OpenRouter, with Reka AI as the only fallback under Decosa's account.","gpus":"operator-contracted","third_parties":[],"retention":"None: the review and the reply live in memory for the request. The signed record holds the review's SHA-256, never its text.","used_for_training":false,"encrypted_while_processed":false},"sealed_tier":{"applies":false,"note":"The sealed tier (raw chat only, never use-case pipelines) is paused at launch (/docs/sealed-tier)."},"external_calls":[]},"console":{"href":"/tools/operations/review-reply","input":"runner","lanes":[{"id":"blocked","title":"What was blocked","kind":"list"},{"id":"reply","title":"Your reply","kind":"markdown"},{"id":"record","title":"Signed record","kind":"json"}],"samples":[{"n":1,"id":"dentist-named","title":"Dentist named","deep_link":"/tools/operations/review-reply?sample=1&autorun=0"},{"n":2,"id":"therapist-angry","title":"Therapist angry","deep_link":"/tools/operations/review-reply?sample=2&autorun=0"},{"n":3,"id":"chiro-mixed","title":"Chiro mixed","deep_link":"/tools/operations/review-reply?sample=3&autorun=0"},{"n":4,"id":"bakery-happy","title":"Bakery happy","deep_link":"/tools/operations/review-reply?sample=4&autorun=0"},{"n":5,"id":"plumber-refund","title":"Plumber refund","deep_link":"/tools/operations/review-reply?sample=5&autorun=0"},{"n":6,"id":"injection","title":"Injection","deep_link":"/tools/operations/review-reply?sample=6&autorun=0"}],"deep_link_params":{"sample":"1-based index into samples, or a sample id","autorun":"1 = start the run once the sample is loaded; 0 (default) = only preselect","reduce-motion":"1 = turn off animations"}},"api":{"base":"https://api.decosa.ai","contract":"/api/contract.json","contract_markdown":"/api/contract.md","reference":"/docs/api","keys":"/account/keys"},"prompts":{"hosted":"/prompts/review-reply-hosted.md","selfhost":"/prompts/review-reply-selfhost.md","assemble":"/prompts/review-reply-assemble.md","mac":null},"rehearsal":{"bundle":"/samples/review-reply.zip","bundle_url":"https://decosa.ai/samples/review-reply.zip","folder":"/samples/review-reply/","expected":"/samples/review-reply/expected.json","files":["/samples/review-reply/expected.json","/samples/review-reply/inputs/review.json"],"bytes":1499,"checks":["the practice is treated as healthcare","the final reply passes every check","the reply never names the dentist","the reply never names the staff member","the reply never uses the reviewer's name","the reply never mentions the crown","the reply is signed off with the practice name","an owner's reply that confirms the visit is refused","the refusal names patient privacy","the signed record verifies","a record with its source changed no longer verifies","every model call has a signed receipt"],"licence":"Synthetic: the practice, people and review are invented. Part of decosa-api.","about":"A synthetic review for a made-up dental practice. The reply must pass every check, never confirm the reviewer is a patient, never name the dentist, the staff member or the reviewer, never mention the crown or the insurance; every model call has a signed receipt; the signed record verifies and a tampered one fails.","run":{"containers":"docker compose exec api python scripts/rehearse.py review-reply","checkout":"python scripts/rehearse.py review-reply --bundle review-reply.zip --base-url http://127.0.0.1:8445","mac":".venv/bin/python scripts/rehearse.py review-reply"},"guidance":"Set up with a coding agent (we recommend Claude Code with Claude Opus 5.5; any capable coding agent works) on mock data only, run the rehearsal until every check passes, then run your own data locally yourself. Never give the agent real data during setup."},"hardware_fit":{"check":"/self-host/hardware?use=review-reply","data":"/api/hardware.json","tiers":[{"id":"lite","gpu_gb":40,"basis":"estimate","unknown":[]},{"id":"standard","gpu_gb":57.6,"basis":"stack","unknown":[]},{"id":"best","gpu_gb":192,"basis":"stack","unknown":[]},{"id":"wanted","gpu_gb":384,"basis":"estimate","unknown":[]}],"mac":null},"links":{"page":"/tools/operations/review-reply","json":"/use-cases/review-reply.json","metrics":"/metrics/review-reply","console":"/tools/operations/review-reply","console_sample":"/tools/operations/review-reply?sample=1&autorun=0","stack":"/tools/operations/review-reply#stack","try_live":"/tools/operations/review-reply","watch":"/tools/operations/review-reply","build":"/tools/operations/review-reply#build","self_host":"/tools/operations/review-reply#self-host","prompts":{"hosted":"/prompts/review-reply-hosted.md","selfhost":"/prompts/review-reply-selfhost.md","assemble":"/prompts/review-reply-assemble.md","mac":null}}}