182 · Healthcare · Sales and marketing · live
Review reply with patient privacy
Eval results
Scored on a held-out or test splitRun 29 Sep 2026Eval write-up (decosa-api, access required)
- Healthcare replies that confirm a patient (blind judge)0 / 40test splitn = 40held-out set generated after the guard was frozen; target 0
- Replies a business could post as written (blind judge)93 / 100test splitn = 100healthcare 38/40, other 55/60
- Healthcare replies that fell back to the fixed safe reply24 / 40test splitn = 40the price of blocking broadly
- Replies with an offer or a fact nobody gave (blind judge)4 / 100test splitn = 100
- First guard on the development set: healthcare replies that confirm a patient9 / 40dev (tuned on)n = 40code rules only; led to the model privacy check and broader rules
Dataset
200 synthetic reviews (80 for health and care businesses) written by Qwen3.8-27B from seeded plans: set 1 for development, set 2 held out and generated after the guard was frozen. Replies judged blind by Claude Code (Opus), which saw only the business type, stars, review and reply.
Caveats
- One judge (a frontier model), no human labels.
- Synthetic reviews written by the same model family that drafts the replies.
- English only.
- The guard blocks broadly, so many health and care replies are the generic fixed reply.
Nightly smoke check
Loading the nightly status…
- Result
- pass
- Run
- 29 Sep 2026
- Latency, this run
- n/a
- p50 over passed runs
- 1.8 s
- Receipts
- 1.55
- Model calls
- n/a
- Tokens
- n/a
- Cost per run
- <$0.001
Self-host verification
Not yet verified on a fresh self-host setup.
Rehearsal bundle: review-reply.zip (1 KB, 12 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.
Known limits
- Timings and cost were measured on the pre-release server through the production gateway. On production (30 Sep 2026) the samples, a made-up review typed in by hand and the own-reply check were run end to end in a browser; self-hosting from the assemble prompt has not been verified yet.
- Reviews are synthetic, written by the same model family that drafts the replies; real reviews are messier.
- English only.
- A new phrasing that confirms a patient can still slip past both checks; read every reply before posting.
Models and licences
Standard tier. Licence posture: permissive (Apache, MIT or BSD).
- The review-reply guard: offers, contact details and numbers not given, placeholders, arguing, and the patient-privacy rules. Deterministic code, the same in TypeScript and Python.@decosa/site-kit review guardApache-2.0
- Drafts the reply (and a redraft when the code checks block the first); for health and care businesses, a second call reads the reply alone and says whether it confirms a patient. The checks themselves are code (the open-source site kit).Qwen3.8-27B (NVIDIA NVFP4)Apache-2.0
All quality evidence
Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.
Lite · one 48 GB card (1)
- accuracy on this task: not measured yet
Standard · the hosted demo, one 96 GB card (2)
- held-out healthcare replies that confirm a patient (blind judge): 0 of 40decosa-api docs/evals/review-reply.md, gateway route, 29 Sep 2026
- held-out replies a business could post as written (blind judge): 93 of 100decosa-api docs/evals/review-reply.md
Best · DeepSeek-V4-Flash on two more cards (1)
- accuracy on this task: not measured yet
Wanted · two large judges from different families (1)
- accuracy on this task: not measured yet