Skip to content
decosa

The full write-up behind the numbers on the tool’s page, as the team that built it wrote it: data, method, results and caveats. Internal names are removed; nothing else is edited.

Eval: Find the songs in your mix (use case 180, mix-cue-sheet)

29 Sep 2026. Engine decosa_api/cue 0.1.0 (ported from cue-engine 1618429). Scorer decosa_api/cue/score.py: a true song start counts as found when some predicted start lies within the tolerance; the first song (0:00) is left out on both sides. CPU only; no model calls in any hosted configuration, so no receipts and no model cost.

(b)+(c) Held-out synthetic mixes (the numbers the site shows)

Data. 34 songs rendered with Decosa Studio's Make a song (ACE-Step 1.5, MIT licence; 60 s each, the tool's cap), through the normal Studio API one job at a time on 29 Sep. Only songs the similarity check called "clear" were used (5 "review" and 1 "near_copy" were left out). Two pools that never mix: d songs for development and the demo mix, t songs (19 clear) for the test set. scripts/cue_synth_mixes.py built 20 test mixes (seed 99): 4-8 songs each, every song played at 0.96-1.04 speed (vinyl style) and trimmed 0-6 s at each end, joined by cuts, equal-power crossfades (4-12 s) or bass-swap blends (12-24 s). The true start is the first sample of the song that is heard. 105 song starts (35 per style), 99 minutes of audio. The test set was built after the engine and every threshold were frozen on the dev split (3 dev mixes and the demo mix, 21 starts; the own-refs refinement, REFINE_MAX_S and the landmark thresholds were set there) and was run once.

Inputs (hosted runs the first three) within 1 s within 5 s within 10 s predicted time per mix p50 / p95
Titles only (the tracklist without times) 15/105 51/105 84/105 105 0.85 / 1.25 s
Your own track files (c) 24/105 92/105 105/105 105 4.46 / 6.26 s
Track files + titles 24/105 92/105 105/105 105 4.62 / 6.08 s
Audio only (novelty, auto count) 3/105 10/105 19/105 20 0.56 / 0.73 s
Titles only, learned detector on (self-host) 11/105 40/105 61/105 105 0.51 / 0.69 s
Track files, detector on (self-host) 14/105 55/105 85/105 105 4.58 / 5.81 s
Audio only, detector on (self-host) 8/105 22/105 30/105 43 0.49 / 0.71 s

Re-run on 30 Sep 2026 after the QA-sweep fixes to track-file matching (stray chance hits of a track file are dropped; a start that several windows place inside the first identifying window is kept). Those rules were set on the tester's own material, not on this test set; the track-file rows above are the re-run (test-eval-2026-09-30.json), the other rows and the timings are the first run. Before the fixes the track-file rows read 23/105, 87/105 and 101/105.

By transition style (within 5 s / within 10 s, of 35 each):

cut crossfade blend
Titles only 26 / 30 11 / 30 14 / 24
Track files 28 / 35 31 / 35 33 / 35

What it means:

  • With the track files, 92 of 105 starts land within 5 s and 105 within 10 s. The misses at 5 s are mostly cuts into songs whose first seconds were skipped: the fingerprint says where the song's first second would be, and the first-heard refinement cannot always tell a quiet intro from the moment of the cut.
  • With titles only, the engine knows how many songs there are and places each at the strongest change; cuts are easy (26/35 within 5 s), long blends hard.
  • The learned detector (trained on real club mixes of 4-7-minute songs) is worse than the novelty segmenter here: these mixes have 60-second songs. It stays off in the hosted tool anyway (licence, page 99 D4).
  • Audio only is not useful on short songs (it finds 20 boundaries for 105 songs): the tool asks for a tracklist or files.

(a) Eight public DJ mixes (internal, development numbers, not held out)

The original engine's development set: 8 public mixes with the DJs' own timestamps, 190 song starts (dancehall, jungle, peak-time house, long melodic-house live sets, progressive, techno, hardstyle). The audio stays on the owner's machine and is never bundled or redistributed. The fingerprint probes were cached from a personal-use recognition service during development; they are not part of the product and no service was called for this eval. Run 29 Sep with the port's numpy features and the detector exported to ONNX.

Configuration 1 s 5 s 10 s 20 s 30 s predicted
Fingerprint probes + titles + learned detector (original engine: 99/123/144) 26 74 99 123 144 277
Same, detector off 13 37 70 95 146 276
Titles only, detector off (the hosted setting) 12 40 61 75 86 187
Titles only, detector on 13 48 57 73 81 187
Audio only, detector on 17 64 78 104 129 358
Audio only, detector off 11 38 59 70 81 176
  • The port reproduces the original exactly (99/123/144) with the detector on: the numpy features match librosa's to 1e-5 (spectral contrast to 0.04 dB) and the ONNX export matches the PyTorch model to 5e-7.
  • Titles only on real long mixes is weak (61/190 within 10 s): long beat-matched blends and 2.5-hour live sets. This is why the page leads with "add the track files you played" and says plainly that titles alone are a rough placement.
  • These mixes were used while building the original engine, so they are development numbers.

Cost and time per task

CPU only, one process, no model calls: $0.00 in model cost for every hosted configuration. Wall clock on our server (shared CPU, 29 Sep): titles only p50 0.85 s / p95 1.25 s per 5-minute mix; with track files p50 4.46 s / p95 6.26 s. The hosted demo end to end over HTTP (6-minute demo mix, 5 runs each on the pre-release server, including the chaptered .m4a and the signed cue sheet): titles p50 6.96 s / p95 7.34 s; with the 8 track files p50 12.01 s / p95 13.25 s. On the 8 public mixes (7-150 minutes) features took about 3.5 s per hour of audio on a Mac Studio CPU.

Expected properties of the sample run (rehearsal/mix-cue-sheet)

  1. The demo mix with its eight track files gives 8 songs in the tracklist's order.
  2. Every start is within 6 s of the true start (recorded run: all 7 within 5 s; the cut to "Slap Happy" at 0.2 s).
  3. Every song is placed from its own track file (source own_refs).
  4. The chaptered .m4a has 8 chapters with the same titles; the tracklist starts "0:00 Mirrorball Sunday".
  5. The signed cue sheet verifies at /record/verify, and fails once a title is edited.
  6. The mix and the track files are deleted when the run ends (inputs_deleted true).

Caveats

  • Synthetic: generated songs of 60 s, mixed by a script; real DJ mixes have longer songs and subtler blends. The same author wrote the mixer and the engine. Each test song appears in about 5 of the 20 mixes (19 songs, 105 slots).
  • Tolerances: DJs' own timestamps disagree by about 9 s between annotators, so "within 10 s" is the practical bar.
  • No frontier comparison: the task is signal processing, not judgment.