98 · Legal · live
Privileged call notes
Eval results
Scored on a held-out or test splitRun 28 Sep 2026
- Memo fact recall (blind grader)89.3%test splitn = 20Dev: 92.2%. Grader: Claude Code Opus 5.5, blind, saw only script, ground truth and memo.
- Memo lines invented / with a wrong detail0 / 12 of 796test splitn = 7967 of the 12 in one call (a year the call never said); the no-model year check now flags them
- Conflicts-name recall, strict / phonetic89.6% / 94.0%test splitn = 20Strict misses are mostly speech-recognition spellings; no invented names
- Dated deadlines found100%test splitn = 20Dates overall: 92.0%
- Recording notice heard, right26/26syntheticn = 26All 26 calls (dev and test). 6 calls never mention recording and are flagged
- Time entry = call length rounded up to 0.1 h26/26syntheticn = 26All 26 calls (dev and test). Narrative acceptable 20/20, client email 14/20 (blind grader)
Dataset
26 synthetic attorney-client calls (6 dev, 20 test) in 9 practice areas, 3 in Spanish, voiced with Decosa house voices (Kokoro-82M) with crosstalk and a telephone band on half; run through the diarizer and the full pipeline.
Caveats
- Synthetic calls voiced by TTS (cleaner than real phone audio); the same model family wrote the scripts and runs the pipeline.
- Fixes after cold-user tests and after the first test runs were mechanisms (spelled names, the client on the list, a weekday and year check), not tuned to test values; the year check was applied post-hoc to the final outputs.
- The claim check blocks invented lines but marked only 7 of 12 lines the grader called wrong (with the year check).
- Time savings are estimates from two blind persona tests, not measured with real lawyers.
Nightly smoke check
Loading the nightly status…
- Result
- pass
- Run
- 28 Sep 2026
- Latency, this run
- n/a
- p50 over passed runs
- 148 s
- Receipts
- 44
- Model calls
- n/a
- Tokens
- n/a
- Cost per run
- $0.017
Self-host verification
Verified on 28 Sep 2026: fresh clone of decosa-api into a clean directory on our server, api image built from docker/api/Dockerfile, compose api service with a named data volume on the host network, direct route, local signing, against the already-running Voxtral, MOSS diarizer and Qwen3.8; torn down after
Assembly prompt §6: TX/CA gives all-party; 400 without consent; custody call upload 39 s, 11 names incl. opposing counsel, 0.1 h, audio dropped, record verified (135 entries), 41 receipts attested. Live replay at 2x: 155 s, 52 captions, record verified (238 entries).
Rehearsal bundle: privileged-call-notes.zip (3 KB, 10 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.
Known limits
- Not a certified transcript: speech recognition mishears names and numbers. Check anything that matters against the call.
- The conflicts list is the names heard on the call; it does not search your conflicts database.
- Deadlines are the ones said on the call. It recomputes the arithmetic, but it does not know court rules or limitation periods.
- The time entry is the call audio's length rounded up to 0.1 hour; your firm's and the client's billing rules decide the entry.
- The hosted demo is for synthetic calls. On the hosted API the server sees the audio and text in memory while it works; real client calls belong on your own box or the Confidential tier.
Models and licences
Standard tier. Licence posture: permissive (Apache, MIT or BSD).
- Live captions during the call (streaming, no speakers)Voxtral Mini 4B RealtimeApache-2.0
- After hang-up (or on an upload): who said what, one line per turn, each with its own receiptMOSS-Transcribe-Diarize 0.9BApache-2.0
- Speaker roles, live intake checklist, cited memo, conflicts names, claim check, time-entry narrative and client emailQwen3.8-27B (NVIDIA NVFP4)Apache-2.0
All quality evidence
Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.
Lite · one 48 GB card, uploaded calls (1)
- call memo quality: not measured yet
Standard · the hosted demo (3)
- memo fact recall, 20 synthetic test calls: 89.3%blind grader, 20 synthetic test calls; decosa-api docs/evals/privileged-call-notes.md (28 Sep 2026)
- conflicts-name recall, 20 synthetic test calls: 89.6% strict / 94.0% phonetic20 synthetic test calls; decosa-api docs/evals/privileged-call-notes.md (28 Sep 2026)
- invented memo items left after the claim check: 0 of 796blind grader on the full memo; 12 items (1.5%) had a wrong detail; decosa-api docs/evals/privileged-call-notes.md (28 Sep 2026)
Best · DeepSeek-V4-Flash writes and checks (1)
- call memo quality: not measured yet