Eval: promotional-claims pre-check (use case 22)
Run 25 Sep 2026 on our server, gateway route, Qwen3.8-27B (NVFP4, vLLM 0.29.0), temperature 0, thinking off. Results:
promo-claims-check-dev.json, promo-claims-check-results.json (test), promo-claims-check-results-repeat.json (the
same test run again, unchanged code). Script: scripts/promo_eval.py; cases: scripts/promo_cases.py.
Data
12 promotional pieces, all synthetic, written for this eval for a fictional distributor ("Fernhill"). References are
public: three FDA labels from openFDA (DailyMed set ids in promo_cases.py; excerpts of the boxed warning,
indications, contraindications, warnings, adverse reactions and clinical studies, 20-23k characters each), three NIH
Office of Dietary Supplements fact sheets for health professionals (public domain; 8k, 11k and 72k characters), and
synthetic one-line product spec sheets for the supplements.
| split | products | pieces | planted (sentence, category) | planted piece-level | unplanted claim sentences |
|---|---|---|---|---|---|
| dev | metformin (Rx), magnesium (supplement) | 4 | 11 | 2 | 31 |
| test | atorvastatin, lisinopril (Rx); vitamin D, omega-3 (supplements) | 8 | 20 | 2 | 63 |
Each product has one piece with planted problems and one written to comply. Planted categories: unsupported (any
support finding: unsupported, contradicted, overstated, or an unbacked comparison or superlative), off_label,
comparative (comparative or superlative without backing), disease (supplement disease claim). Piece level:
fair_balance (missing, thin, or no boxed warning) and dshea (no disclaimer). Every sentence that is not planted
should raise no issue-level finding. That is the false-positive measure.
The cases were written and committed (f3cbb04) before the first model run. The test split was never used to change anything. One change was made after the first dev run: supplement comparatives are backed when the references state the comparison, and the head-to-head check is kept for Rx drugs and devices. That change came from a dev false positive ("citrate has higher bioavailability than magnesium oxide" was flagged for lacking a head-to-head trial). Nothing else was tuned. The prompts are as first written.
Results
| dev | test | test, repeat | |
|---|---|---|---|
| planted problems caught (all categories) | 11 / 11 | 20 / 20 | 20 / 20 |
| unsupported / off-label / comparative / disease | 4/4, 2/2, 2/2, 3/3 | 9/9, 3/3, 3/3, 5/5 | 9/9, 3/3, 3/3, 5/5 |
| piece-level (fair balance, DSHEA) | 2 / 2 | 2 / 2 | 2 / 2 |
| unplanted sentences with an issue | 2 / 31 | 1 / 63 (1.6%) | 1 / 63 |
| unplanted sentences with any finding (issue or check) | 2 / 31 | 3 / 63 | 4 / 63 |
| piece-level findings on pieces with none planted | 0 | 0 | 0 |
| compliant pieces with no issue | 2 / 2 | 3 / 4 | 3 / 4 |
The one test false positive, in both runs: "In ASCOT, atorvastatin 10 mg reduced coronary events by 36% in patients with hypertension and multiple risk factors" was marked off-label. The label's indication covers adults with multiple risk factors for CHD; the reviewer read "hypertension" as a population outside it. The check-level findings on compliant copy were "overstated" on a headline ("blood pressure control you can count on") and on a safety line that gave one upper limit where the fact sheet gives a range by age.
In dev, two issue-level findings landed on unplanted sentences of the planted magnesium page. One was the headline "the magnesium your body is missing", called a disease claim; a reviewer could argue it either way. The other was "supports healthy muscle and nerve function", called unsupported because the fact sheet backs magnesium, not the named product.
Run-to-run variation. Greedy decoding on a shared, batched server is not bit-for-bit repeatable. While recording the site's replay fixture (four demo samples, twice), the compliant metformin piece once came back with one issue. The judge misread the label's flattened Table 8 (the openFDA text lists the columns out of order), called the glyburide-arm number contradicted, and so the head-to-head comparison became unbacked. In the other recording and in both dev runs the same sentence was supported. Tables flattened to text are a known weak spot.
Cost and speed (test, 8 pieces): 156 model calls, all receipted; 13.0k generated and 677k prompt tokens (about 1.6k generated per piece; prompts are large because every call resends the reference pack, which prefix caching absorbs). A piece took 25-77 s end to end through the shared gateway, on a GPU shared with other work.
What this does and does not show
- The planted problems are blatant: they are the kind a first-pass reviewer should catch, written by the same person who built the checker. 20/20 means it reliably catches obvious problems. It does not show that it catches subtle ones: a misleading-but-true statistic, an omitted qualifier, or a claim that is fine in isolation but misleading in context. An eval on real OPDP untitled and warning letters (public, with the cited violations) is the next step.
- The false-positive rate is on copy written to comply, 94 unplanted sentences in all. On real copy, with brand voice, headlines and testimonials, expect more check-level noise.
- Text only. Visual prominence of risk information, which OPDP letters often cite, is not measured.
- Devices were not evaluated (no public device labeling in this set).