Skip to content
decosa

182 · Healthcare · Sales and marketing · live

Review reply with patient privacy

Open the toolJSON

Eval results

Scored on a held-out or test splitRun 29 Sep 2026Eval write-up (decosa-api, access required)

  • Healthcare replies that confirm a patient (blind judge)0 / 40test splitn = 40held-out set generated after the guard was frozen; target 0
  • Replies a business could post as written (blind judge)93 / 100test splitn = 100healthcare 38/40, other 55/60
  • Healthcare replies that fell back to the fixed safe reply24 / 40test splitn = 40the price of blocking broadly
  • Replies with an offer or a fact nobody gave (blind judge)4 / 100test splitn = 100
  • First guard on the development set: healthcare replies that confirm a patient9 / 40dev (tuned on)n = 40code rules only; led to the model privacy check and broader rules

Dataset

200 synthetic reviews (80 for health and care businesses) written by Qwen3.8-27B from seeded plans: set 1 for development, set 2 held out and generated after the guard was frozen. Replies judged blind by Claude Code (Opus), which saw only the business type, stars, review and reply.

Caveats

  • One judge (a frontier model), no human labels.
  • Synthetic reviews written by the same model family that drafts the replies.
  • English only.
  • The guard blocks broadly, so many health and care replies are the generic fixed reply.

Nightly smoke check

Loading the nightly status…

Result
pass
Run
29 Sep 2026
Latency, this run
n/a
p50 over passed runs
1.8 s
Receipts
1.55
Model calls
n/a
Tokens
n/a
Cost per run
<$0.001

Self-host verification

Not yet verified on a fresh self-host setup.

Rehearsal bundle: review-reply.zip (1 KB, 12 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.

Known limits

  • Timings and cost were measured on the pre-release server through the production gateway. On production (30 Sep 2026) the samples, a made-up review typed in by hand and the own-reply check were run end to end in a browser; self-hosting from the assemble prompt has not been verified yet.
  • Reviews are synthetic, written by the same model family that drafts the replies; real reviews are messier.
  • English only.
  • A new phrasing that confirms a patient can still slip past both checks; read every reply before posting.

Models and licences

Standard tier. Licence posture: permissive (Apache, MIT or BSD).

  • The review-reply guard: offers, contact details and numbers not given, placeholders, arguing, and the patient-privacy rules. Deterministic code, the same in TypeScript and Python.@decosa/site-kit review guardApache-2.0
  • Drafts the reply (and a redraft when the code checks block the first); for health and care businesses, a second call reads the reply alone and says whether it confirms a patient. The checks themselves are code (the open-source site kit).Qwen3.8-27B (NVIDIA NVFP4)Apache-2.0

All quality evidence

Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.

Lite · one 48 GB card (1)
  • accuracy on this task: not measured yet
Standard · the hosted demo, one 96 GB card (2)
  • held-out healthcare replies that confirm a patient (blind judge): 0 of 40decosa-api docs/evals/review-reply.md, gateway route, 29 Sep 2026
  • held-out replies a business could post as written (blind judge): 93 of 100decosa-api docs/evals/review-reply.md
Best · DeepSeek-V4-Flash on two more cards (1)
  • accuracy on this task: not measured yet
Wanted · two large judges from different families (1)
  • accuracy on this task: not measured yet

How we measure · All tools