Skip to content
decosa

Developers · Building blocks

Song starts in long audio

Where each song starts in a DJ mix, a radio show or any long recording. It uses what you know first: times in your tracklist, then your own track files found in the audio, then the audio's own change points for the rest, marked estimated. CPU only, no model call, and nothing leaves the server.

Measured 2026-09-29. Full eval. First use: Find the songs in your mix.

Watch a real run

Loading the tool…

How it works

  1. ffmpeg decodes the audio once; the engine computes spectral features and a novelty curve of where the music changes (numpy).
  2. A tracklist is parsed line by line. Times you give are kept as written; with times on some lines only, the rest are placed between them.
  3. Your own track files are found in the mix by the landmark fingerprinter from the sample clearance block, searched across tempo changes of up to 6% and the pitch shift that comes with them.
  4. The resolver turns those sightings and the change points into one start per song, in tracklist order. Starts from your times or files are marked identified; the rest are estimated.
  5. ffmpeg writes a chaptered .m4a (AAC copied as it is, other formats encoded once), and the server signs a decosa.record.v1 cue sheet: the audio's SHA-256, the engine version and method, and one entry per song with its start, title and how it was found.

Limits

Mixup to 150 MB and 2 hours, uploaded in chunks of at most 16 MB
Track filesup to 30, each up to 40 MB and 15 minutes
Tracklistup to 400 lines and 20,000 characters; "12:30 Artist - Title" or just "Artist - Title"
ComputeCPU only: no GPU and no model call
Retention (hosted)mix and track files deleted when the run ends; chaptered .m4a kept one hour; song list and cue sheet 24 hours

Speed

6-minute demo mix, titles only (8 songs)7.0 s p50, 7.3 s p95 over 5 runs; the engine itself about 1.6 s, the rest writing the chaptered file
The same mix with its 8 track files12.0 s p50, 13.3 s p95 over 5 runs; finding the files about 5 s
Held-out set, per 5-minute mix (engine)with track files p50 4.46 s, p95 6.26 s; titles only p50 0.85 s, p95 1.25 s

Results

Held out, with your track files: starts within 1 / 5 / 10 s24/105 / 92/105 / 105/105
Held out, titles only (the hosted default)15/105 / 51/105 / 84/105
Held out, audio only (20 starts predicted: not useful on its own)3/105 / 10/105 / 19/105
Held out, by join with track files (within 5 s / 10 s, of 35 each)cuts 23 / 31, crossfades 31 / 35, blends 33 / 35
Held out, by join with titles onlycuts 26 / 30, crossfades 11 / 30, blends 14 / 24
Real DJ mixes (dev), titles only, detector off: within 10 / 20 / 30 s61 / 75 / 86 of 190
Real DJ mixes (dev), with a fingerprint service's identifications, titles and the detector99 / 123 / 144 of 190 (within 5 s: 74); detector off: 70 / 95 / 146

Held out: 20 synthetic mixes of 19 of Decosa's own 60-second songs (105 starts, 99 minutes), joined by cuts, crossfades and bass-swap blends, 35 of each; thresholds frozen on a dev split first, test run once; the same author wrote the mixer and the engine. Dev: 8 public DJ mixes, 190 starts, not held out; long blends make them harder. DJs' own timestamps disagree by about 9 s, so 10 s is the practical bar. The fingerprint identifications came from a service that is not part of the product.

API

POST /cue/uploads{kind: mix | ref, bytes, filename} -> {upload_id, chunk_bytes}; then PUT /cue/uploads/{id}?offset=N with at most 16 MB of raw bytes per chunk (409 names the expected offset).
POST /cue/mark{audio: upload_id | sample: id, tracklist?, refs?: [{upload, title}], title?, stream?} -> SSE run, stage (features, refs, resolve, chapters), done, or 202 {run_id, poll}. Token or dk_ key for mix-cue-sheet.
GET /cue/runs/{id}markers [{start, title, source, confidence}], method, duration_s, audio_sha256, timing and the download paths: tracklist.txt, chapters.m4a (one hour), record.
POST /record/verify{record} -> ok, summary, bad: checks the signed cue sheet. No key.
decosa-api: decosa_api.cuepipeline.mark(path, tracklist=None, refs=None) returns the markers, method and timing; chapters.write_m4a and chapters.tracklist_text write the files. Apache-2.0, usable as a library.

Where it goes next

  • Find the songs in your mixbuilt

    Drop in a DJ mix: every song start, a tracklist for the upload, a chaptered .m4a and a signed cue sheet.

  • The Cue iPhone appunblocked

    Its player skips a chaptered .m4a song by song, so a mix you mark here plays like an album.

  • Radio and podcast archivesunblocked

    Segment long shows from a running order and the station's own music files.

  • Music video and Studio editsupgrade

    Cut points for a long set from where the songs change.

Licences

decosa-cue engine (decosa_api/cue)Apache-2.0 (ported from an MIT-licensed original); open-source package publishing soon
Landmark fingerprinter (decosa_api/verticals/clearance)AGPL-3.0-or-later
numpyBSD-3-Clause
ffmpeg (run as a separate program)LGPL-2.1+ / GPL-2.0+
Transition detector v1 (optional, self-host only)weights not released; training audio (DJ Mix Dataset, YouTube-sourced) has no dataset licence
Demo mixeight songs made with Decosa Studio's Make a song (ACE-Step 1.5, MIT), mixed by Decosa

What it does not do

  • Naming songs nobody gave it: there is no catalogue lookup on the hosted API. A self-hosted server can turn on an AudD lookup with its own token, which sends the audio to AudD.
  • The learned transition detector on the hosted API: it stays off until it is retrained on licence-clean audio.
  • Speed beyond the 6-minute demo mix: longer mixes are not timed yet.
All building blocks