07 · Public sector · Compliance and trust · live
Tamper-evident record
Eval results
Not held outRun 23 Sep 2026
- Claim check on the synthetic sessions: sentences supported14/17 council, 9/9 interviewsyntheticn = 26Re-run 26 Sep 2026 after the sessions were re-voiced with Decosa house voices (Kokoro-82M). 2 of the 3 council flags come from one summary sentence split at "Mr."; the first build (macOS voices): 19/19 and 9/9.
- Speaker attribution on the synthetic council meeting (5 TTS voices)5 speakers found; one short turn given to the wrong membersyntheticn = 5Re-voiced 26 Sep 2026 with Decosa house voices (Kokoro-82M) and re-run; the first build (macOS voices): 4 speakers found, two male voices merged.
Dataset
Two synthetic TTS sessions (a council meeting and an interview) replayed through the live stack on a decosa-api test instance, one session at a time, gateway route.
Caveats
- WER on meeting or interview audio not measured yet; the ASR numbers published elsewhere are on clinical consultations (PriMock57).
- Summary quality on meetings not measured yet.
- Verifier recall on meeting minutes not measured yet.
- Two synthetic sessions only.
Nightly smoke check
Loading the nightly status…
- Result
- pass
- Run
- 25 Sep 2026
- Latency, this run
- n/a
- p50 over passed runs
- 17 s
- Receipts
- 54
- Model calls
- n/a
- Tokens
- n/a
- Cost per run
- $0.009
Self-host verification
Verified on 25 Sep 2026: fresh clone, api image built, the prompt's .env and compose used as written, sample against local model servers
The step 6 smoke passed as written: the record verified (132 entries), and after one edited word verification failed at the first edited transcript line. With the diarize service the record also carried speaker labels. Verified on 2026-09-25: images build, services start, sample passes end to end against local model servers equivalent to the documented ones; model-server startup itself not re-verified.
Rehearsal bundle: record.zip (630 KB, 9 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.
Known limits
- Proves the record was not changed after signing and which key signed it; it does not prove the speech was recognised correctly. Not a certified court record.
- Speaker labels need the diarize service; without it transcript lines have no speaker names.
Models and licences
Standard tier. Licence posture: permissive (Apache, MIT or BSD).
- Live captions (streaming, no speakers)Voxtral Mini 4B RealtimeApache-2.0
- After the session: speaker-attributed transcript, one line per turnMOSS-Transcribe-Diarize 0.9BApache-2.0
- Actions lane, speaker roles, cited summary, claim verifierQwen3.8-27B (NVIDIA NVFP4)Apache-2.0
All quality evidence
Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.
Lite · one 48 GB card, captions only (1)
- WER on meeting audio: not measured yet
Standard · the hosted demo, two passes (3)
- MOSS-TD WER / DER on PriMock57 (clinical proxy): 10.3 / 11.4scribe-bench RESULTS.md
- claim check on the synthetic sessions: sentences supported: 19/19 and 9/9measured on our server 2026-09-23: a decosa-api pre-release test instance, scripts/replay_client.py, one session at a time, gateway route
- WER on meeting audio: not measured yet
Best · DeepSeek-V4-Flash writes and checks (1)
- summary quality on meetings: not measured yet