Skip to content
decosa

49 · Film, TV and games · Sales and marketing · live

Synthetic-performer disclosure and S&P pre-flight

Open the toolJSON

Eval results

Scored on a held-out or test splitRun 25 Sep 2026Eval write-up (decosa-api, access required)

  • Planted script issues found (full config)27 of 29 (93%)test splitn = 29Dev: 31 of 32
  • Precision of script flags (full config)1.00 (27 flags, 0 false)test splitn = 27
  • Clean scripts with a flag0 of 8test splitn = 8Without the yes/no step: 2 of 8; word lists only: 2 of 8
  • Word lists only (no model): planted issues found7 of 29test splitn = 29
  • OCR reads the AI label on finished Decosa ads / false alarm on unlabelled shots18 of 18 frames / 0 of 60 framessyntheticn = 78Own label style on own footage only

Dataset

48 synthetic ad scripts for fictional brands (24 dev, 24 test), with issues planted from fixed pools that share nothing between dev and test; Decosa's own UGC ad renders and ten raw H3 shots for the label check.

Caveats

  • The plants are blatant, one clear instance each, in short synthetic scripts; expect lower recall and precision on real scripts.
  • The same author wrote the eval script, the planted pools and the prompts; one prompt change was made after the first dev run.
  • The label check measures Decosa's own label style on its own footage; third-party labels may not be read.
  • Conspicuousness under NY GBL 396-b is a legal judgement and is not measured. No synthetic-performer detector is used or claimed.

Nightly smoke check

Loading the nightly status…

Result
pass
Run
25 Sep 2026
Latency, this run
n/a
p50 over passed runs
3.6 s
Receipts
5
Model calls
n/a
Tokens
n/a
Cost per run
$0.001

Self-host verification

Verified on 25 Sep 2026: Fresh clone into a clean directory, docker build (50 s), the api service with named volumes, a development C2PA certificate from scripts/provenance_devcert.py, pointed at the running local vLLM (Qwen3.8-27B) over host networking; then torn down.

Verified on 2026-09-25: the image has ffmpeg, Tesseract 5.5.0, the font and c2pa-python; the planted script gives the same five flags and consent states as the hosted run in 2.7 s (five attested calls); the raw clip gets the label (6 of 6 frames read back) and a valid C2PA marking in 5.3 s; the attestation verifies and fails when one decision is changed; no script text or names in the logs. The Decosa-ad samples are hosted renders and show as unavailable, as expected. The model server's own startup was not re-verified (no new GPU load).

Rehearsal bundle: disclosure-preflight.zip (254 KB, 16 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.

Known limits

  • Advisory, not a legal clearance; 'conspicuous' is not measured.
  • Synthetic performers are found only from provenance or a declaration.
  • Script recall is measured on short synthetic scripts with blatant plants.
  • Logos and faces in the picture are not checked yet.
  • Consent comes from the consent ledger (tool 47) by id; the hosted demo's id_demo identities are fictional. A server without the ledger reports real performers as unchecked, never cleared.

Models and licences

Standard tier. Licence posture: permissive (Apache, MIT or BSD).

  • The pre-flight: provenance lookup, performer evidence, rules for NY and the EU, consent lookups by ledger id, the profanity and music-cue word lists, the report, sign-off and the signed attestation (no model; CPU)decosa-api disclosure module (decosa_api/verticals/disclosure)AGPL-3.0-or-later
  • Reads six frames for an AI label (whole frame, then the top and bottom bands), before and after the label is addedTesseract OCR 5 (English)Apache-2.0
  • Reads the file's C2PA credential and signs the marking: the source as a parentOf ingredient, c2pa.opened and c2pa.edited actions with the IPTC digital source type, and an ai.decosa.disclosure assertionc2pa-python 0.37 (c2pa-rs)MIT OR Apache-2.0
  • Decodes Decosa's invisible video watermark (TrustMark), so a Decosa render whose credential was stripped is still recognisedTrustMark Q (decoder)MIT
  • Model: lists candidate names, brands, profanity, rating triggers and music cues in the script, then answers a typed yes/no for each (typed-judgment, calibrated probability)Qwen3.8-27B (NVFP4)Apache-2.0

All quality evidence

Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.

How well does it catch planted issues, and does the label stick?

  • Planted script issues found (test, 29): 93% (27) (27 flags, none false; 0 of 8 clean scripts flagged)
  • Without the typed yes/no step: 28 found, 2 false (both false flags were generic music cues)
  • Word lists only (no model): 7 of 29 (profanity and music markup only)
  • Label read back after burn-in: 60 of 60 frames (and 0 of 60 frames of unlabelled clips read as labelled)
  • C2PA marking valid, content intact: 10 of 10

Source: decosa-api docs/evals/disclosure-preflight.md, 2026-09-25

Lite · code only, any CPU (3)
  • Planted script issues found / precision / clean scripts flagged (test, 29 plants in 24 synthetic scripts, 8 clean): 7 of 29 / 0.78 / 2 of 8decosa-api docs/evals/disclosure-preflight.md, 2026-09-25 (word lists only: profanity 4 of 4, music markup 3 of 6; the false flags are generic music cues)
  • AI label read on Decosa ads / read back after burn-in / false label on unlabelled clips: 18 of 18 / 60 of 60 / 0 of 60 framesdecosa-api docs/evals/disclosure-preflight.md, 2026-09-25 (3 finished ads, 10 raw MiniMax H3 shots)
  • C2PA marking valid with content intact and the disclosure assertion: 10 of 10decosa-api docs/evals/disclosure-preflight.md, 2026-09-25
Standard · adds script checks by Qwen3.8-27B (hosted demo) (4)
  • Planted script issues found / precision / clean scripts flagged (test): 27 of 29 (93%) / 1.00 (27 flags, 0 false) / 0 of 8decosa-api docs/evals/disclosure-preflight.md, 2026-09-25 (dev run twice, one change to the rating question between runs; test run once)
  • By category (test): real person / brand / profanity / rating trigger / music: 4 of 4 / 7 of 8 / 4 of 4 / 7 of 7 / 5 of 6decosa-api docs/evals/disclosure-preflight.md, 2026-09-25 (misses: 'Apple' and the song 'Happy')
  • Without the yes/no step (test): found / precision / clean scripts flagged: 28 of 29 / 0.93 / 2 of 8decosa-api docs/evals/disclosure-preflight.md, 2026-09-25
  • Demo samples giving the expected performers, rules, consent states and flags: 7 of 7decosa-api docs/evals/disclosure-preflight.md, 2026-09-25

How we measure · All tools