38 · Music · Legal · live
Sample and lyric clearance pre-check
Eval results
Scored on a held-out or test splitRun 25 Sep 2026Eval write-up (decosa-api, access required)
- Audio plants found, medium and high confidence62 of 108 (57%)test splitn = 108
- Audio precision, medium and high confidence0.87 (9 false items)test split
- Audio plants found, high confidence only (precision)35 of 108 (32%), precision 1.00test splitn = 108
- Clean tracks with any flag3 of 30test splitn = 30
- Lyric plants found with the model's labels: exact / near / heavy paraphrase8/8, 12/12, 1/8test splitn = 281 false item; 0 of 15 clean songs flagged.
- Chromaprint baseline, samples found4 of 84 (616 chance matches)test splitn = 84
Dataset
Audio: 35 CC BY 4.0 Kevin MacLeod reference recordings and 73 host recordings by the same artist, with synthetic plants (mixed slices under pitch, tempo, filter, loop and varispeed transforms, plus re-played melody interpolations) and clean excerpts; dev and test use different hosts and seeds. Lyrics: 156 US public-domain songs (5,435 lines) planted into 60 model-written songs, 30 dev and 30 test.
Caveats
- Samples are mixed synthetically; real productions add compression, reverb and other layers. No real-world recall is claimed.
- Thresholds were tuned on dev; the test split was scored twice, before and after an interpolation rule changed on dev (first run: 62 of 108, precision 0.89).
- One artist's music throughout; a catalog of many artists may behave differently.
- Only references in the catalog can be found; the hosted catalog is a 41-reference demo.
- Heavy lyric paraphrases are mostly out of reach; the lyric set is English and pre-1929.
Nightly smoke check
Loading the nightly status…
- Result
- pass
- Run
- 25 Sep 2026
- Latency, this run
- n/a
- p50 over passed runs
- 6.7 s
- Receipts
- 1
- Model calls
- n/a
- Tokens
- n/a
- Cost per run
- <$0.001
Self-host verification
Verified on 25 Sep 2026: Fresh clone into a clean directory, docker build, the api service with a named volume, embedding files and the demo catalog fetched inside the container, pointed at the running local vLLM (Qwen3.8-27B) over host networking; then torn down.
Verified on 2026-09-25: the image builds with ffmpeg and the clearance extra, the demo catalog builds in 81 s (download included), the planted sample returns the same five items and the declared-not-found entry as the hosted run in 5.4 s on the direct route (one attested call), the clean sample returns nothing, the signed record verifies and fails when one action is changed, and no lyric or title text reaches the logs. The first attempt found a bug (the demo track list was not in the image), fixed in the branch. The model server's own startup was not re-verified (no new GPU load).
Rehearsal bundle: sample-clearance.zip (649 KB, 10 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.
Known limits
- Recall is moderate: 57% of planted borrowings on the test split; quiet samples and short loops are missed most.
- It only finds what is in the catalog it is given; the hosted demo catalog is 41 references.
- The eval mixes are synthetic, from one composer's CC BY music; real productions may behave differently.
- Lyric matching is English and needs a lyric set: the demo uses public-domain songs published before 1929.
Models and licences
Standard tier. Licence posture: permissive (Apache, MIT or BSD).
- Audio matcher: log-frequency landmarks with pitch and tempo search, peak-by-peak verification, melody shingles for composition references; lyric spans; declared vs detected; the signed record (no model; CPU)decosa-api clearance module (decosa_api/verticals/clearance)AGPL-3.0-or-later
- Lyric-line embeddings: re-worded lines that share few characters with the originalall-MiniLM-L6-v2 (ONNX)Apache-2.0
- Model: labels each near-duplicate lyric line lift, variant, stock phrase or differentQwen3.8-27B (NVFP4)Apache-2.0
All quality evidence
Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.
How well does it find planted borrowings?
- Samples and melodies found (108 plants): 57% (62) (precision 0.87; 3 of 30 clean tracks got a medium-confidence flag)
- High-confidence items: 35 found, all correct (0 of 30 clean tracks flagged)
- Samples mixed at -3 / -6 / -9 / -12 dB: 71% / 57% / 57% / 43%
- Lyric lines: exact / small changes / heavy paraphrase: 8/8 · 12/12 · 1/8 (no clean song flagged)
- Chromaprint alone, same samples: 4 of 84 (and 616 chance matches: built for whole recordings)
Source: decosa-api docs/evals/sample-clearance.md, 2026-09-25
Lite · code only, any CPU (4)
- Planted samples and melodies found / precision / clean tracks flagged (test split, 108 plants in CC BY music, 30 clean tracks): 62 of 108 (57%) / 0.87 / 3 of 30decosa-api docs/evals/sample-clearance.md, 2026-09-25 (thresholds set on the dev split; test scored twice, before and after one dev change, both reported)
- High-confidence items only: found / precision / clean tracks flagged: 35 of 108 / 1.00 / 0 of 30decosa-api docs/evals/sample-clearance.md, 2026-09-25
- By transform (test): pitch ±1-2 semitones / tempo ±5-10% / low-pass / high-pass / 2 s loops / interpolations: 14 of 24 / 16 of 24 / 5 of 6 / 6 of 6 / 2 of 6 / 12 of 24decosa-api docs/evals/sample-clearance.md, 2026-09-25
- Lyric lines, code only (test): exact / one or two words changed / heavy paraphrase / clean songs flagged: 8 of 8 / 12 of 12 / 0 of 8 / 0 of 15decosa-api docs/evals/sample-clearance.md, 2026-09-25 (with embeddings; without them, trigrams alone)
Standard · adds lyric labels by Qwen3.8-27B (hosted demo) (2)
- Lyric lines with the model's labels (test): exact / one or two words changed / heavy paraphrase / false items / clean songs flagged: 8 of 8 / 12 of 12 / 1 of 8 / 1 / 0 of 15decosa-api docs/evals/sample-clearance.md, 2026-09-25; 12 model calls, 515 generated tokens for 30 songs
- Audio (same matcher as Lite): 62 of 108 found, precision 0.87decosa-api docs/evals/sample-clearance.md, 2026-09-25