Eval: sample and lyric clearance pre-check (38)
Run on our server, 25 Sep 2026. Full numbers: docs/evals/sample-clearance/results.json. Scripts: scripts/clearance_eval.py
(audio), scripts/clearance_lyrics_eval.py (lyrics), scripts/clearance_plant.py, scripts/clearance_melody.py.
Data and licences
| What | Source | Licence |
|---|---|---|
| Reference catalog: 35 recordings | Kevin MacLeod, incompetech.com (list: decosa_api/verticals/clearance/data/demo_tracks.json, with each Wikimedia Commons page) |
CC BY 4.0 |
| Host tracks: 73 other recordings by the same artist, not in the catalog | same | CC BY 4.0 |
| 6 melodies (composition references) and their re-played plants | generated here (clearance_melody.py, seeds 38001-38006) |
ours, CC0 |
| Lyric reference set: 156 songs, 5,435 lines | English Wikisource, Category:Song_lyrics, pages giving a year of 1928 or earlier | public domain in the US |
| 60 synthetic songs that receive the lyric plants | written by Qwen3.8-27B through the gateway on 60 neutral topics (receipted) | ours |
No copyrighted material was used. Using one artist for catalog and hosts makes the task harder, not easier: the same instruments and production recur across tracks, so chance resemblance is more likely than across a real catalog.
Audio
Plants. Each planted track is a 75 s excerpt of a host with one slice of a catalog recording mixed in at -3, -6, -9 or -12 dB relative to the host, then encoded to MP3 128 kb/s. Slices are 3-10 s (loops: 1-2.5 s repeated 4 times). Transforms, cycled: none; pitch +1, -1, +2, -2 semitones (Rubber Band, tempo kept); tempo 0.9, 0.95, 1.05, 1.1 (pitch kept); low-pass 1.5 kHz; high-pass 500 Hz; loop x4; varispeed 0.95 and 1.05 (pitch and tempo together). Interpolations: eight or more bars of a catalog melody re-played on another instrument (flute, square or organ), 1-3 semitones up or down, tempo 0.88-1.12, at 0 or -3 dB over a host. Clean tracks: host excerpts with nothing added.
Splits. Dev: the first third of the hosts, seed 1000 (84 samples, 12 interpolations, 30 clean). Test: the other hosts, seed 2000 (84 samples, 24 interpolations, 30 clean). A found item must name the right reference and overlap the planted span (1 s slack). A false item is any flagged reference that was not planted there.
Tuning (dev only). Landmark thresholds: at least 20 dense votes and at least 0.20 of the reference's peaks found above background (a grid over votes 9-40 and verification 0.10-0.40). Interpolations: two agreeing landmark segments on the melody, or a melody run whose best alignment stands 0.45 above the reference's median (or 0.35 over 8 windows with one landmark segment). The test split was scored twice: once before the interpolation rule was changed on dev and once after. Both runs are reported.
Results (test)
| Found | Precision | Clean tracks with any flag | |
|---|---|---|---|
| All items (medium and high) | 62 of 108 (57%) | 0.87 (9 false items) | 3 of 30 |
| High confidence only | 35 of 108 (32%) | 1.00 | 0 of 30 |
| First test run (before the interpolation change) | 62 of 108 | 0.89 (8 false) | 3 of 30 |
| Chromaprint alone (the brief's suggestion, as a baseline) | 4 of 84 samples | 616 chance matches |
By transform (test, medium and high): none 3/6, pitch 14/24, tempo 16/24, low-pass 5/6, high-pass 6/6, loop 2/6, varispeed 4/12, interpolation 12/24. By mix level (samples): -3 dB 20/28, -6 dB 16/28, -9 dB 8/14, -12 dB 6/14. Dev for comparison: 58 of 96 (60%), precision 0.91, 0 of 30 clean tracks flagged.
What the false items are: 3 of 9 flag "Meditation Impromptu 02" inside "Meditation Impromptu 01" and "03", a series that may share material; 3 are interpolation flags of synthetic melodies over busy hosts; 3 are single sample flags just above the threshold. All were medium confidence.
What it misses: quiet slices (-9 and -12 dB) and short loops, where too few reference peaks survive the mix; varispeed (pitch and tempo shifted together by a fraction of a semitone, between the grid points); some plain copies in dense passages. Chromaprint is built to name whole recordings and does not survive mixing, which is why landmarks lead.
Speed. One 75 s track: 2-6 s with four threads while the machine was loaded (11 s once, under heavier load); the eval ran 16 tracks at once on a shared 32-core box at a median of 25 s per track. The demo catalog index (41 references) builds in about 10 s and is 2.8 MB.
Lyrics
Plants. Half of the 30 songs per split get 1-3 plants from the public-domain set: an exact line (6-14 words); the same line with one or two words changed by a fixed word list ("near"); or the line heavily re-worded by the model, "keep its image, change most of the words" ("variant"). The other 15 songs stay clean. Dev: songs 1-30, test 31-60. Thresholds (5-word spans, trigram cosine 0.55, embedding cosine 0.70) were set on dev.
| Test | Exact | Near | Heavy paraphrase | False items | Clean songs flagged | Model calls |
|---|---|---|---|---|---|---|
| With the model's labels | 8/8 | 12/12 | 1/8 | 1 | 0 of 15 | 12 (515 generated tokens) |
| Code only (no model) | 8/8 | 12/12 | 0/8 | 0 | 0 of 15 | 0 |
Heavy paraphrases ("Was loved by a blushing rose" -> "Adored by a crimson bloom") are mostly out of reach: they share meaning, not words, and flagging on meaning alone flags every love song. The model's labels add one paraphrase and one false item on this set; their real job is dismissing stock phrases ("I love you so"), which the demo shows.
Limits
- The catalog is the ceiling: nothing outside it can be found, and the hosted catalog is a 41-reference demo.
- Samples are mixed synthetically; real productions add compression, reverb and other layers. No real-world recall is claimed.
- One artist's music throughout; a catalog of many artists may behave differently (probably fewer false items).
- The lyric set is English and pre-1929; modern lyrics are copyrighted and must come from the customer.