Eval: synthetic-performer disclosure and S&P pre-flight (49)
Run on our server on 25 Sep 2026, through the model gateway while other workloads were also using it. The full numbers are in
docs/evals/disclosure-preflight/, and the script is scripts/disclosure_eval.py.
Data and licences
| What | Source | Licence |
|---|---|---|
| 48 base ad scripts for fictional brands | Written by Qwen3.8-27B through the gateway: one receipted call each, seed 4900+i, ids in bases.json |
ours |
| Planted issues | Fixed pools in the eval script: 24 public figures, 24 brands, 12 swear words, 12 rating lines, 12 song titles, 12 hard negatives | n/a (names only) |
| Video | Decosa's own renders: three finished UGC ads and ten raw MiniMax H3 shots from the same briefs | ours |
I read every base script by hand for accidental real names, brands, songs or rating content. I found none and excluded none. The bases only mention their own fictional brand and websites like "Orla.com".
Splits. Dev and test are alternate bases, 24 each. The test pools share no person, brand, word, stunt, song or hard negative with dev. Each split has:
- 8 clean scripts, each carrying two hard negatives: "this deal is a steal", "killing time", "[MUSIC: soft original guitar]", a fictional "GRANDMA ROSE", and so on;
- 16 planted scripts, each with 1 to 3 issues and sometimes one hard negative.
That gives 32 plants in dev and 29 in test.
Scoring. A plant counts as found when a flag of the same category sits on its line or names it. A flag counts as false when it matches no plant.
- Precision is computed over flags of confidence "high" and "review".
- A music cue with no existing song named is an "info" note, not a flag, and is left out of both counts.
Tuning. Dev was run twice. After the first run I made one change: the rating question now lists every rating
category, not only the subtype the extraction guessed. That took dev rating recall from 4/6 to 5/6. The test split was
run once, after that change. KEEP_P = 0.5 and HIGH_P = 0.8 were set before any run and never changed.
Script flags
| Config | Split | Found | Precision (flags) | Clean scripts with a flag |
|---|---|---|---|---|
| Full: word lists + Qwen extraction + typed yes/no | test | 27 of 29 (93%) | 1.00 (27 flags, 0 false) | 0 of 8 |
| Full | dev | 31 of 32 | 1.00 (31, 0 false) | 0 of 8 |
| No yes/no step | test | 28 of 29 | 0.93 (2 false) | 2 of 8 |
| Word lists only (no model) | test | 7 of 29 | 0.78 (2 false) | 2 of 8 |
Test results by category, full config:
| Category | Found |
|---|---|
| Real person | 4/4 |
| Brand | 7/8 |
| Profanity | 4/4 |
| Rating trigger | 7/7 |
| Music | 5/6 |
The two misses are single words that are also everyday words: "Apple" in "I cancelled Apple for this", and the song "Happy".
Every false flag without the yes/no step was a generic music cue ("[MUSIC: light library piano track]"). The yes/no step turns those into info notes. It also dropped one true song cue on test, "Happy", which is the price of that step. The extraction itself raised no false flag on either split.
Cost and speed (test, full). 50 receipted calls for 24 scripts. That is about 1,000 prompt tokens and 100 generated tokens a script, roughly $0.0005 at list price. The median time was 1.3 s a script with at most two calls in flight.
What this does not show. The plants are blatant, one clear instance each, in short synthetic scripts. Real scripts are messier:
- nicknames and partial names;
- brands used as generic words;
- music described without a title;
- rating context that depends on the picture.
Expect lower recall and precision on real material. Profanity comes from a word list, so a swear word the list lacks is found only by the model.
Visible label and marking
| Check | Result |
|---|---|
| OCR reads the label on the three finished Decosa ads | 18 of 18 sampled frames |
| OCR reads a label on the ten raw, unlabelled H3 shots (false alarm) | 0 of 60 frames |
| Label burned into the ten raw shots, then read back | 60 of 60 frames (10 of 10 clips on every frame) |
C2PA marking on those ten: valid, content intact, ai.decosa.disclosure present |
10 of 10 |
The OCR is Tesseract 5.3.4. It reads the whole frame first, then the top and bottom quarters at 2x in greyscale. The band pass was added after one frame of a busy shot failed the whole-frame pass during development.
All ten shots are 1080x1920 H3 renders, so this measures our own label style on our own footage. A third party's label, in another font, position or language, may not be read. The check then says "missing" or "partial", and the reviewer decides.
Conspicuousness is not measured. The rule counts a label as visible when OCR reads it on every sampled frame. Whether that is "conspicuous" under NY GBL § 396-b is a legal judgement.
Demo samples
All seven samples gave the expected results on the pre-release server against the gateway:
- performer kinds;
- the NY, EU 50(2) and EU 50(4) statuses;
- consent states (granted, revoked, purpose not covered);
- flags;
- whether anything was applied.
The stripped-credential sample shows the watermark path: the Decosa receipt is found by the invisible watermark alone, and the marking is signed again.
Synthetic-performer detection
No detector model is used or claimed. A performer is synthetic when:
- the file's provenance says so: a Decosa render receipt found by hash, credential or watermark, or another generator's C2PA credential declaring AI media; or
- the uploader declares it.
A file with neither gets "needs declaration". This is by design: a detector miss would read as "no AI here".
Verdict
Would a buyer pay?
- For Decosa's own UGC ads: yes. It is the missing last step. The ad ships with a checked label, a signed marking and an attestation naming the reviewer. For an AI-ad studio selling into New York, that is a line item.
- For outside uploads: maybe. The value depends on the uploader declaring performers honestly, and on real performers being enrolled in the consent ledger.
- For network S&P desks: not yet. The text half works on planted scripts. The picture half (logos, faces, on-screen text in frames) is not built, and the wiki's S&P interviews have not happened.
What is missing:
- The picture half. The UGC OCR brand check is the licence-clean starting point.
- Third-party label styles in the OCR check.
- An eval on real scripts with real clearance notes.
- Real enrolled performers depend on the consent ledger (use case 47), which this branch now includes and checks by id
(tested). The hosted demo still uses fictional
id_demoidentities in a private ledger.