22 · Healthcare · Compliance and trust · live
Promotional-claims pre-check
Eval results
Scored on a held-out or test splitRun 25 Sep 2026Eval write-up (decosa-api, access required)
- Planted problems caught (all categories)20 / 20test splitn = 20Repeat run of the same test split: 20 / 20. Dev: 11 / 11.
- Piece-level problems caught (fair balance, DSHEA)2 / 2test splitn = 2
- Unplanted sentences with an issue (false positives)1 / 63 (1.6%)test splitn = 63Repeat run: 1 / 63. Dev: 2 / 31.
- Unplanted sentences with any finding (issue or check)3 / 63test splitn = 63Repeat run: 4 / 63.
- Compliant pieces with no issue3 / 4test splitn = 4
Dataset
12 synthetic promotional pieces for a fictional distributor (4 dev, 8 test), checked against public FDA labels (openFDA), NIH Office of Dietary Supplements fact sheets and synthetic spec sheets. Each product has one piece with planted problems and one written to comply.
Caveats
- The planted problems are blatant and were written by the same person who built the checker; subtle problems are not measured.
- Synthetic copy only; the false-positive rate is on copy written to comply, and real copy should bring more check-level noise.
- Text only: visual prominence of risk information is not measured, and devices were not evaluated.
- Greedy decoding on a shared server is not bit-for-bit repeatable; one recording run flagged a compliant piece.
- One change was made after the first dev run; the test split was never used to change anything.
Nightly smoke check
Loading the nightly status…
- Result
- pass
- Run
- 25 Sep 2026
- Latency, this run
- n/a
- p50 over passed runs
- 3.9 s
- Receipts
- 14
- Model calls
- n/a
- Tokens
- n/a
- Cost per run
- $0.022
Self-host verification
Verified on 25 Sep 2026: fresh clone, compose up, sample against local model servers
Method: a fresh clone of decosa-api main, the api image built from it, the compose file from this prompt, then the prompt's smoke steps and the nightly smoke module, against the already-running local Qwen3.8-27B vLLM. Verified on 25 Sep 2026: images build, services start, sample passes end to end against local model servers equivalent to the documented ones; model-server startup itself not re-verified. Planted metformin piece: status issues with every expected finding (contradicted HbA1c, off-label weight and heart claims, comparative, fair balance, boxed warning); the compliant piece came back clean; the packet verifies and a changed finding fails.
Rehearsal bundle: promo-claims-check.zip (5 KB, 10 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.
Known limits
- Speed depends on load: a 7-14 sentence piece took 4-10 s on a quiet GPU and 30-35 s while the shared GPU was busy (25 Sep 2026).
- Text only: type size, placement and contrast are not assessed. Tables flattened to text can be misread.
- The eval's planted problems are blatant ones; subtle violations are not measured yet. A person reviews every finding.
Models and licences
Standard tier. Licence posture: permissive (Apache, MIT or BSD).
- Pre-check: sentences, label sections, rules, findings, the packet (no model; CPU)decosa-api promo module (decosa_api/verticals/promo) with the grounding module (decosa_api/verticals/grounding)AGPL-3.0-or-later
- Grounding judge, claim reviewer, head-to-head and fair-balance checksQwen3.8-27B (NVFP4)Apache-2.0
All quality evidence
Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.
Standard · one GPU for the model (hosted demo) (4)
- Planted problems caught, held-out test (unsupported, off-label, comparative, disease): 20 / 20 (9/9, 3/3, 3/3, 5/5), same in a repeat rundocs/evals/promo-claims-check.md: 8 synthetic pieces on 3 public labels and 3 NIH fact sheets
- Piece-level problems caught, test (fair balance, DSHEA disclaimer): 2 / 2docs/evals/promo-claims-check.md
- False positives, test: unplanted sentences with an issue / with any finding: 1 / 63 (1.6%) / 3-4 of 63docs/evals/promo-claims-check.md, two runs
- Dev split (the only split used for a change): 11 / 11 caught; 2 / 31 unplanted with an issuedocs/evals/promo-claims-check.md