Skip to content
decosa

07 · Public sector · Compliance and trust · live

Tamper-evident record

Open the toolJSON

Eval results

Not held outRun 23 Sep 2026

  • Claim check on the synthetic sessions: sentences supported14/17 council, 9/9 interviewsyntheticn = 26Re-run 26 Sep 2026 after the sessions were re-voiced with Decosa house voices (Kokoro-82M). 2 of the 3 council flags come from one summary sentence split at "Mr."; the first build (macOS voices): 19/19 and 9/9.
  • Speaker attribution on the synthetic council meeting (5 TTS voices)5 speakers found; one short turn given to the wrong membersyntheticn = 5Re-voiced 26 Sep 2026 with Decosa house voices (Kokoro-82M) and re-run; the first build (macOS voices): 4 speakers found, two male voices merged.

Dataset

Two synthetic TTS sessions (a council meeting and an interview) replayed through the live stack on a decosa-api test instance, one session at a time, gateway route.

Caveats

  • WER on meeting or interview audio not measured yet; the ASR numbers published elsewhere are on clinical consultations (PriMock57).
  • Summary quality on meetings not measured yet.
  • Verifier recall on meeting minutes not measured yet.
  • Two synthetic sessions only.

Nightly smoke check

Loading the nightly status…

Result
pass
Run
25 Sep 2026
Latency, this run
n/a
p50 over passed runs
17 s
Receipts
54
Model calls
n/a
Tokens
n/a
Cost per run
$0.009

Self-host verification

Verified on 25 Sep 2026: fresh clone, api image built, the prompt's .env and compose used as written, sample against local model servers

The step 6 smoke passed as written: the record verified (132 entries), and after one edited word verification failed at the first edited transcript line. With the diarize service the record also carried speaker labels. Verified on 2026-09-25: images build, services start, sample passes end to end against local model servers equivalent to the documented ones; model-server startup itself not re-verified.

Rehearsal bundle: record.zip (630 KB, 9 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.

Known limits

  • Proves the record was not changed after signing and which key signed it; it does not prove the speech was recognised correctly. Not a certified court record.
  • Speaker labels need the diarize service; without it transcript lines have no speaker names.

Models and licences

Standard tier. Licence posture: permissive (Apache, MIT or BSD).

All quality evidence

Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.

Lite · one 48 GB card, captions only (1)
  • WER on meeting audio: not measured yet
Standard · the hosted demo, two passes (3)
  • MOSS-TD WER / DER on PriMock57 (clinical proxy): 10.3 / 11.4scribe-bench RESULTS.md
  • claim check on the synthetic sessions: sentences supported: 19/19 and 9/9measured on our server 2026-09-23: a decosa-api pre-release test instance, scripts/replay_client.py, one session at a time, gateway route
  • WER on meeting audio: not measured yet
Best · DeepSeek-V4-Flash writes and checks (1)
  • summary quality on meetings: not measured yet

How we measure · All tools