05 · Field and trades · live
Field reports
Eval results
Not held outRun 23 Sep 2026
- Expected issues in the report (grounded report + safety sweep + claim check)40/41 (98%); before 38/41 (93%)syntheticn = 41
- Severity within the accepted range, of issues found40/40; all 6 critical-only hazards rated critical (before 35/38 and 6/6)syntheticn = 40
- Stated facts kept51/51; before 32/51 (63%)syntheticn = 51incl. passed checks and spec limits
- Claims the judge flagged as contradicted or unsupported5/125syntheticn = 1252 are the overall-condition rating; 3 are speech-recognition spellings quoted verbatim. By hand: no invented findings.
Dataset
Internal synthetic eval: 8 TTS inspection walk-throughs (2 recorded demo + 6 new) with a gold checklist and a Gemma 4 judge, run through the live stack on decosa-api b24ab2a. The TTS audio was macOS system voices; this eval was not re-run after the demo audio was re-voiced with Decosa house voices (Kokoro-82M) on 26 Sep 2026.
Caveats
- The prompts were tuned on these same 8 scripts, so there is no held-out set.
- Synthetic TTS audio only; no real site recordings.
- Small n (8 walk-throughs).
- Model judge (Gemma 4), checked by hand only in part.
Nightly smoke check
Loading the nightly status…
- Result
- pass
- Run
- 25 Sep 2026
- Latency, this run
- n/a
- p50 over passed runs
- 26 s
- Receipts
- 28
- Model calls
- n/a
- Tokens
- n/a
- Cost per run
- $0.026
Self-host verification
Verified on 25 Sep 2026: fresh clone, api image built, the prompt's .env and compose used as written, sample against local model servers
The step 5 replay passed as written (32 attested receipts; checklist, issues, measurements, passed_checks, then safety_sweep, report_check and report), and step 6 exported the report to JSON, Markdown and PDF (10 issues: 9 supported, 1 partial). Verified on 2026-09-25: images build, services start, sample passes end to end against local model servers equivalent to the documented ones; model-server startup itself not re-verified.
Rehearsal bundle: field.zip (665 KB, 9 checks). Mock inputs plus the expected results, so you can prove your own setup works before any real data touches it.
Known limits
- The report is written after the walk-through ends: about 25 s on the hosted demo for a 75 s sample, longer under load.
- Measurements are checked against the spec the inspector states; there is no built-in code database.
Models and licences
Standard tier. Licence posture: permissive (Apache, MIT or BSD).
- Speech recognition (streaming)Voxtral Mini 4B RealtimeApache-2.0
- Lanes and report (language model)Qwen3.8-27B (NVIDIA NVFP4)Apache-2.0
All quality evidence
Every sourced number on the tool’s Stack tab, by tier. Some are proxies from another task; their labels say so.
Lite · 4-bit on a 32 GB card, plus a small card for speech (3)
- Owner's coding benchmark, core tasks (/500), Qwen3.8-27B NVFP4 + MTP on vLLM: 478coding-agent-bench README (our server)
- Same weights served by llama.cpp (Q4 GGUF path): avoid for this tier: 62–63 / 500coding-agent-bench README (our server)
- Field lane quality: not measured yet
Standard · one 96 GB card (the hosted demo) (8)
- Owner's coding benchmark, core tasks (/500), Qwen3.8-27B NVFP4 + MTP on vLLM: 478coding-agent-bench README (our server)
- ROUGE-L on ACI-Bench clinical notes (proxy, not a field task; FP8 on vLLM): 34.2scribe-bench wiki/models.md
- Quality smoke, math / code: 38/40, 12/12Decosa model benchmarks (Sep 2026)
- Field report, internal synthetic eval (n = 8 walk-throughs: 2 recorded demo + 6 new), with the grounded report, safety sweep and claim check: expected issues in the report: 40/41 (98%); before 38/41 (93%)eval results file lanes-field-grounded-synthetic.json (internal synthetic eval: the same 8 TTS walk-throughs, gold checklist and Gemma 4 judge as the 23 Sep 2026 eval, re-run on decosa-api b24ab2a; the prompts were tuned on these same 8 scripts, so there is no held-out set; 2026-09-23)
- Same eval: severity within the accepted range, of issues found: 40/40; all 6 critical-only hazards rated critical (before 35/38 and 6/6)eval results file lanes-field-grounded-synthetic.json (internal synthetic eval: the same 8 TTS walk-throughs, gold checklist and Gemma 4 judge as the 23 Sep 2026 eval, re-run on decosa-api b24ab2a; the prompts were tuned on these same 8 scripts, so there is no held-out set; 2026-09-23)
- Same eval: stated facts kept (incl. passed checks and spec limits, now fields of their own): 51/51; before 32/51 (63%)eval results file lanes-field-grounded-synthetic.json (internal synthetic eval: the same 8 TTS walk-throughs, gold checklist and Gemma 4 judge as the 23 Sep 2026 eval, re-run on decosa-api b24ab2a; the prompts were tuned on these same 8 scripts, so there is no held-out set; 2026-09-23)
- Same eval: missed safety item (bathroom receptacles without GFCI): in the report in 3 of 3 re-runs. Run on the original report, the safety sweep adds it as critical, quoted at 1:03eval results file lanes-field-grounded-synthetic.json (internal synthetic eval: the same 8 TTS walk-throughs, gold checklist and Gemma 4 judge as the 23 Sep 2026 eval, re-run on decosa-api b24ab2a; the prompts were tuned on these same 8 scripts, so there is no held-out set; 2026-09-23)
- Same eval: claims the judge flagged as contradicted or unsupported: 5/125 (before 0/100, plus 1 found by hand: 'cracked chimney' where the inspector said sealant). By hand: no invented findings. 2 of the 5 are the report's overall-condition rating; 3 are speech-recognition spellings now quoted verbatim ('ceiling' for sealant, 'pole' for pull stations, 'stab block' for Stab-Lok), which need an ASR glossary, not a model changeeval results file lanes-field-grounded-synthetic.json (internal synthetic eval: the same 8 TTS walk-throughs, gold checklist and Gemma 4 judge as the 23 Sep 2026 eval, re-run on decosa-api b24ab2a; the prompts were tuned on these same 8 scripts, so there is no held-out set; 2026-09-23)
Best · two 96 GB cards (4)
- ROUGE-L on ACI-Bench clinical notes (proxy, not a field task; run on a Mac via MLX, not the NVFP4 kit): 35.8scribe-bench wiki/models.md
- Medical-term miss rate on PriMock57, MOSS-Transcribe-Diarize vs Nemotron-3.5 streaming (%): 8.4 vs 12.7scribe-bench RESULTS.md / wiki/decoder-finding.md
- Owner's coding benchmark, core tasks (/500), DeepSeek V4-Flash: 453–459coding-agent-bench README (our server)
- Field lane quality: not measured yet
Wanted · the largest open flash models (1)
- Field lane quality: not measured yet