Skip to content
decosa

The full write-up behind the numbers on the tool’s page, as the team that built it wrote it: data, method, results and caveats. Internal names are removed; nothing else is edited.

Eval: synthetic-performer disclosure and S&P pre-flight (49)

Run on our server on 25 Sep 2026, through the model gateway while other workloads were also using it. The full numbers are in docs/evals/disclosure-preflight/, and the script is scripts/disclosure_eval.py.

Data and licences

What Source Licence
48 base ad scripts for fictional brands Written by Qwen3.8-27B through the gateway: one receipted call each, seed 4900+i, ids in bases.json ours
Planted issues Fixed pools in the eval script: 24 public figures, 24 brands, 12 swear words, 12 rating lines, 12 song titles, 12 hard negatives n/a (names only)
Video Decosa's own renders: three finished UGC ads and ten raw MiniMax H3 shots from the same briefs ours

I read every base script by hand for accidental real names, brands, songs or rating content. I found none and excluded none. The bases only mention their own fictional brand and websites like "Orla.com".

Splits. Dev and test are alternate bases, 24 each. The test pools share no person, brand, word, stunt, song or hard negative with dev. Each split has:

  • 8 clean scripts, each carrying two hard negatives: "this deal is a steal", "killing time", "[MUSIC: soft original guitar]", a fictional "GRANDMA ROSE", and so on;
  • 16 planted scripts, each with 1 to 3 issues and sometimes one hard negative.

That gives 32 plants in dev and 29 in test.

Scoring. A plant counts as found when a flag of the same category sits on its line or names it. A flag counts as false when it matches no plant.

  • Precision is computed over flags of confidence "high" and "review".
  • A music cue with no existing song named is an "info" note, not a flag, and is left out of both counts.

Tuning. Dev was run twice. After the first run I made one change: the rating question now lists every rating category, not only the subtype the extraction guessed. That took dev rating recall from 4/6 to 5/6. The test split was run once, after that change. KEEP_P = 0.5 and HIGH_P = 0.8 were set before any run and never changed.

Script flags

Config Split Found Precision (flags) Clean scripts with a flag
Full: word lists + Qwen extraction + typed yes/no test 27 of 29 (93%) 1.00 (27 flags, 0 false) 0 of 8
Full dev 31 of 32 1.00 (31, 0 false) 0 of 8
No yes/no step test 28 of 29 0.93 (2 false) 2 of 8
Word lists only (no model) test 7 of 29 0.78 (2 false) 2 of 8

Test results by category, full config:

Category Found
Real person 4/4
Brand 7/8
Profanity 4/4
Rating trigger 7/7
Music 5/6

The two misses are single words that are also everyday words: "Apple" in "I cancelled Apple for this", and the song "Happy".

Every false flag without the yes/no step was a generic music cue ("[MUSIC: light library piano track]"). The yes/no step turns those into info notes. It also dropped one true song cue on test, "Happy", which is the price of that step. The extraction itself raised no false flag on either split.

Cost and speed (test, full). 50 receipted calls for 24 scripts. That is about 1,000 prompt tokens and 100 generated tokens a script, roughly $0.0005 at list price. The median time was 1.3 s a script with at most two calls in flight.

What this does not show. The plants are blatant, one clear instance each, in short synthetic scripts. Real scripts are messier:

  • nicknames and partial names;
  • brands used as generic words;
  • music described without a title;
  • rating context that depends on the picture.

Expect lower recall and precision on real material. Profanity comes from a word list, so a swear word the list lacks is found only by the model.

Visible label and marking

Check Result
OCR reads the label on the three finished Decosa ads 18 of 18 sampled frames
OCR reads a label on the ten raw, unlabelled H3 shots (false alarm) 0 of 60 frames
Label burned into the ten raw shots, then read back 60 of 60 frames (10 of 10 clips on every frame)
C2PA marking on those ten: valid, content intact, ai.decosa.disclosure present 10 of 10

The OCR is Tesseract 5.3.4. It reads the whole frame first, then the top and bottom quarters at 2x in greyscale. The band pass was added after one frame of a busy shot failed the whole-frame pass during development.

All ten shots are 1080x1920 H3 renders, so this measures our own label style on our own footage. A third party's label, in another font, position or language, may not be read. The check then says "missing" or "partial", and the reviewer decides.

Conspicuousness is not measured. The rule counts a label as visible when OCR reads it on every sampled frame. Whether that is "conspicuous" under NY GBL § 396-b is a legal judgement.

Demo samples

All seven samples gave the expected results on the pre-release server against the gateway:

  • performer kinds;
  • the NY, EU 50(2) and EU 50(4) statuses;
  • consent states (granted, revoked, purpose not covered);
  • flags;
  • whether anything was applied.

The stripped-credential sample shows the watermark path: the Decosa receipt is found by the invisible watermark alone, and the marking is signed again.

Synthetic-performer detection

No detector model is used or claimed. A performer is synthetic when:

  • the file's provenance says so: a Decosa render receipt found by hash, credential or watermark, or another generator's C2PA credential declaring AI media; or
  • the uploader declares it.

A file with neither gets "needs declaration". This is by design: a detector miss would read as "no AI here".

Verdict

Would a buyer pay?

  • For Decosa's own UGC ads: yes. It is the missing last step. The ad ships with a checked label, a signed marking and an attestation naming the reviewer. For an AI-ad studio selling into New York, that is a line item.
  • For outside uploads: maybe. The value depends on the uploader declaring performers honestly, and on real performers being enrolled in the consent ledger.
  • For network S&P desks: not yet. The text half works on planted scripts. The picture half (logos, faces, on-screen text in frames) is not built, and the wiki's S&P interviews have not happened.

What is missing:

  • The picture half. The UGC OCR brand check is the licence-clean starting point.
  • Third-party label styles in the OCR check.
  • An eval on real scripts with real clearance notes.
  • Real enrolled performers depend on the consent ledger (use case 47), which this branch now includes and checks by id (tested). The hosted demo still uses fictional id_demo identities in a private ledger.